Multi-Algorithm Optimization of Explanation Stability in Ensemble Learning for Human Capital Attrition Risk Prediction
Attrition models are tuned for discrimination, but the decision they support is a ranking of retention drivers, and nothing requires that ranking to survive retraining. We treat its reproducibility as an objective: hyperparameter optimization for an ensemble of gradient-boosted and bagged learners becomes a bi-objective program maximizing the cost-sensitive area under the precision–recall curve jointly with the expected agreement of SHapley Additive exPlanations (SHAP) global rankings across bootstrap refits. Four results follow, three specializing standard theory. Hoeffding’s U-statistic theory makes the estimator unbiased with variance 4ζ1/R+O(R−2), turning “how many refits” into a variance calculation; an elementary spacing argument makes a ranking’s survival a signal-to-noise rather than a noise-magnitude quantity; SHAP linearity makes the retraining variance in the ensemble attributions a quadratic form, so the mixing weights solve a minimum-variance portfolio. The fourth delimits the third: a ranking is invariant to attribution scale, and this variance is not, so attribution variance is no proxy for explanation stability—minimizing it selects the least stable learner on every benchmark—and stability is optimized directly. The cost of the scheme is given in closed form. Our central empirical finding is a caution. Across five public benchmarks spanning 1.4–50.6% positives, no test separates the multi-objective solvers from the single-objective ones at a practitioner’s budget, and the objective must be estimated to be optimized: at the refit budget such a search would choose, the estimator’s standard deviation exceeds the spread of true stability across the configurations it ranks, so maximizing the estimate overfits it. Re-estimating each search’s own selection with draws that took no part in selecting it leaves a resolvable stability gain on two of the five benchmarks. Giving the search five and 12.5 times the budget surfaces gains that survive where the practitioner’s budget surfaces none, so part of what a cheap search misses is real; but the gap between what a search advertises and what survives does not close as the budget grows, and on one benchmark widens. Declaring the objective is worth doing; what it pays cannot be read off the search that declared it. The caution applies to any model selection that resamples to estimate its criterion and then maximizes it.
Authors
- Fan Si
- Zihe Qi
- Ziwei Chen
Institutions
- University of Illinois Urbana-Champaign (US)
- Peking University (CN)
- Beijing Jiaotong University (CN)
Publication Details
- Journal
- Mathematics
- Published
- 2026-09-10
- DOI
- https://doi.org/10.3390/math14183283
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00