A leave‐top‐feature‐out framework for validating feature importance reliability in concrete compressive strength modeling
Abstract Civil Engineering Design has seen a surge in AI‐driven analyses of concrete compressive strength, yet this proliferation has obscured a methodological crisis in supervised machine learning for sustainable materials research. Existing studies conflate two accuracy dimensions: target prediction accuracy, validatable against known labels, and feature importance accuracy, which lacks validation benchmarks. Establishing true variable associations requires two criteria from causal inference: consistency (stable feature rankings across perturbations) and dose–response relationships (monotonic changes in model output upon feature removal), criteria prior studies have failed to meet. We introduce a leave‐top‐feature‐out validation framework applied to a concrete compressive strength dataset, wherein top‐ranked features are systematically removed and perturbations quantified. Evaluation of supervised models (Random Forest, XGBoost), unsupervised methods (Feature Agglomeration, Highly Variable Gene Selection), and nonparametric statistics (Spearman correlation) reveals that supervised approaches produce volatile rankings susceptible to label‐driven biases, while unsupervised methods yield more consistent hierarchies, with SHapley Additive exPlanations (SHAP) explanations shown to amplify model biases. Our framework establishes the first theoretically grounded benchmark for distinguishing genuine from spurious associations in concrete strength prediction.
Authors
- Yoshiyasu Takefuji (ORCID: https://orcid.org/0000-0002-1826-742X)
Institutions
- NOK Corporation (Japan) (JP)
Publication Details
- Journal
- Civil Engineering Design
- Published
- 2026-09-09
- DOI
- https://doi.org/10.1002/cend.70026
- Primary Topic
- Innovative concrete reinforcement materials
- Type
- article
- Field-Weighted Citation Impact
- 0.00