Predictive accuracy does not establish feature-importance robustness: a leave-top feature-out analysis in late-life depression

Supervised machine learning models are increasingly applied in biomedical signal processing and control for feature assessment; however, a critical epistemological gap persists in that no established ground truth exists against which feature importance can be validated. This paper theoretically establishes that supervised models exhibit two distinct dimensions of accuracy, target prediction accuracy and feature importance accuracy, which are not interchangeable. High target prediction accuracy does not guarantee reliable feature importance, as importance metrics reflect contributions to prediction rather than true causal associations. Furthermore, SHAP-based explanations inherit and may amplify the underlying model’s distortions rather than correct them, rendering their outcomes contingent on model-specific artifacts rather than confirmed causal insights. To address this gap, we draw on the extended Bradford Hill criteria to propose consistency and dose–response relationships as principled validation criteria, and we introduce a novel leave-top-feature-out diagnostic framework to empirically test the stability of feature importance rankings. Empirical analysis of a late-life depression dataset comparing nine feature selection algorithms demonstrates that unsupervised models yield superior stability in feature ranking, whereas supervised models exhibit label-driven instability that SHAP integration fails to resolve, underscoring the need for validation frameworks that extend beyond model-dependent interpretability methods.

Authors

Institutions

Publication Details

Journal
Biomedical Signal Processing and Control
Published
2026-09-24
DOI
https://doi.org/10.1016/j.bspc.2026.111577
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Predictive accuracy does not establish feature-importance robustness: a leave-top feature-out analysis in late-life depression

Yoshiyasu Takefuji
Biomedical Signal Processing and Control
Explainable Artificial Intelligence (XAI)
article

Predictive accuracy does not establish feature-importance robustness: a leave-top feature-out analysis in late-life depression

Yoshiyasu Takefuji
article en

Abstract

Supervised machine learning models are increasingly applied in biomedical signal processing and control for feature assessment; however, a critical epistemological gap persists in that no established ground truth exists against which feature importance can be validated. This paper theoretically establishes that supervised models exhibit two distinct dimensions of accuracy, target prediction accuracy and feature importance accuracy, which are not interchangeable. High target prediction accuracy does not guarantee reliable feature importance, as importance metrics reflect contributions to prediction rather than true causal associations. Furthermore, SHAP-based explanations inherit and may amplify the underlying model’s distortions rather than correct them, rendering their outcomes contingent on model-specific artifacts rather than confirmed causal insights. To address this gap, we draw on the extended Bradford Hill criteria to propose consistency and dose–response relationships as principled validation criteria, and we introduce a novel leave-top-feature-out diagnostic framework to empirically test the stability of feature importance rankings. Empirical analysis of a late-life depression dataset comparing nine feature selection algorithms demonstrates that unsupervised models yield superior stability in feature ranking, whereas supervised models exhibit label-driven instability that SHAP integration fails to resolve, underscoring the need for validation frameworks that extend beyond model-dependent interpretability methods.

Biomedical Signal Processing and ControlVol. 129
SciencePark Corporation (Japan) (JP)
Openalex Percentile: Top 9%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Predictive accuracy does not establish feature-importance robustness: a leave-top feature-out analysis in late-life depression — Yoshiyasu Takefuji · Biomedical Signal Processing and Control (2026) | TGRS Research Map | TGRS