Predictive accuracy does not establish feature-importance robustness: a leave-top feature-out analysis in late-life depression
Supervised machine learning models are increasingly applied in biomedical signal processing and control for feature assessment; however, a critical epistemological gap persists in that no established ground truth exists against which feature importance can be validated. This paper theoretically establishes that supervised models exhibit two distinct dimensions of accuracy, target prediction accuracy and feature importance accuracy, which are not interchangeable. High target prediction accuracy does not guarantee reliable feature importance, as importance metrics reflect contributions to prediction rather than true causal associations. Furthermore, SHAP-based explanations inherit and may amplify the underlying model’s distortions rather than correct them, rendering their outcomes contingent on model-specific artifacts rather than confirmed causal insights. To address this gap, we draw on the extended Bradford Hill criteria to propose consistency and dose–response relationships as principled validation criteria, and we introduce a novel leave-top-feature-out diagnostic framework to empirically test the stability of feature importance rankings. Empirical analysis of a late-life depression dataset comparing nine feature selection algorithms demonstrates that unsupervised models yield superior stability in feature ranking, whereas supervised models exhibit label-driven instability that SHAP integration fails to resolve, underscoring the need for validation frameworks that extend beyond model-dependent interpretability methods.
Authors
- Yoshiyasu Takefuji (ORCID: https://orcid.org/0000-0002-1826-742X)
Institutions
- SciencePark Corporation (Japan) (JP)
Publication Details
- Journal
- Biomedical Signal Processing and Control
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1016/j.bspc.2026.111577
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00