Multi-Algorithm Optimization of Explanation Stability in Ensemble Learning for Human Capital Attrition Risk Prediction

Attrition models are tuned for discrimination, but the decision they support is a ranking of retention drivers, and nothing requires that ranking to survive retraining. We treat its reproducibility as an objective: hyperparameter optimization for an ensemble of gradient-boosted and bagged learners becomes a bi-objective program maximizing the cost-sensitive area under the precision–recall curve jointly with the expected agreement of SHapley Additive exPlanations (SHAP) global rankings across bootstrap refits. Four results follow, three specializing standard theory. Hoeffding’s U-statistic theory makes the estimator unbiased with variance 4ζ1/R+O(R−2), turning “how many refits” into a variance calculation; an elementary spacing argument makes a ranking’s survival a signal-to-noise rather than a noise-magnitude quantity; SHAP linearity makes the retraining variance in the ensemble attributions a quadratic form, so the mixing weights solve a minimum-variance portfolio. The fourth delimits the third: a ranking is invariant to attribution scale, and this variance is not, so attribution variance is no proxy for explanation stability—minimizing it selects the least stable learner on every benchmark—and stability is optimized directly. The cost of the scheme is given in closed form. Our central empirical finding is a caution. Across five public benchmarks spanning 1.4–50.6% positives, no test separates the multi-objective solvers from the single-objective ones at a practitioner’s budget, and the objective must be estimated to be optimized: at the refit budget such a search would choose, the estimator’s standard deviation exceeds the spread of true stability across the configurations it ranks, so maximizing the estimate overfits it. Re-estimating each search’s own selection with draws that took no part in selecting it leaves a resolvable stability gain on two of the five benchmarks. Giving the search five and 12.5 times the budget surfaces gains that survive where the practitioner’s budget surfaces none, so part of what a cheap search misses is real; but the gap between what a search advertises and what survives does not close as the budget grows, and on one benchmark widens. Declaring the objective is worth doing; what it pays cannot be read off the search that declared it. The caution applies to any model selection that resamples to estimate its criterion and then maximizes it.

Authors

Institutions

Publication Details

Journal
Mathematics
Published
2026-09-10
DOI
https://doi.org/10.3390/math14183283
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multi-Algorithm Optimization of Explanation Stability in Ensemble Learning for Human Capital Attrition Risk Prediction

Fan Si, Zihe Qi, Ziwei Chen
Mathematics
Explainable Artificial Intelligence (XAI)
article

Multi-Algorithm Optimization of Explanation Stability in Ensemble Learning for Human Capital Attrition Risk Prediction

Fan Si, Zihe Qi, Ziwei Chen
article en

Abstract

Attrition models are tuned for discrimination, but the decision they support is a ranking of retention drivers, and nothing requires that ranking to survive retraining. We treat its reproducibility as an objective: hyperparameter optimization for an ensemble of gradient-boosted and bagged learners becomes a bi-objective program maximizing the cost-sensitive area under the precision–recall curve jointly with the expected agreement of SHapley Additive exPlanations (SHAP) global rankings across bootstrap refits. Four results follow, three specializing standard theory. Hoeffding’s U-statistic theory makes the estimator unbiased with variance 4ζ1/R+O(R−2), turning “how many refits” into a variance calculation; an elementary spacing argument makes a ranking’s survival a signal-to-noise rather than a noise-magnitude quantity; SHAP linearity makes the retraining variance in the ensemble attributions a quadratic form, so the mixing weights solve a minimum-variance portfolio. The fourth delimits the third: a ranking is invariant to attribution scale, and this variance is not, so attribution variance is no proxy for explanation stability—minimizing it selects the least stable learner on every benchmark—and stability is optimized directly. The cost of the scheme is given in closed form. Our central empirical finding is a caution. Across five public benchmarks spanning 1.4–50.6% positives, no test separates the multi-objective solvers from the single-objective ones at a practitioner’s budget, and the objective must be estimated to be optimized: at the refit budget such a search would choose, the estimator’s standard deviation exceeds the spread of true stability across the configurations it ranks, so maximizing the estimate overfits it. Re-estimating each search’s own selection with draws that took no part in selecting it leaves a resolvable stability gain on two of the five benchmarks. Giving the search five and 12.5 times the budget surfaces gains that survive where the practitioner’s budget surfaces none, so part of what a cheap search misses is real; but the gap between what a search advertises and what survives does not close as the budget grows, and on one benchmark widens. Declaring the objective is worth doing; what it pays cannot be read off the search that declared it. The caution applies to any model selection that resamples to estimate its criterion and then maximizes it.

MathematicsVol. 14(18)
University of Illinois Urbana-Champaign (US), Peking University (CN), Beijing Jiaotong University (CN)
Reduced inequalities
Openalex Percentile: Top 8%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.