Interpretable Machine Learning for In-Hospital Mortality Prediction in ICU Patients Using First-24-Hour Routine Vital Signs: A SHAP-Based MIMIC-IV Study
Background: Identification of ICU patients at increased risk of death may support clinical assessment and prioritization of care. However, many mortality prediction approaches incorporate extensive laboratory information, conventional severity scores, or computationally complex models that may limit straightforward interpretation at the bedside. This study investigated whether routinely recorded physiological measurements collected during the first 24 h of ICU care could support an interpretable machine learning approach to subsequent in-hospital mortality prediction. Methods: We performed a retrospective analysis of 51,598 adult first ICU admissions from MIMIC-IV using a strict 24 h landmark design. Patients with a missing ICU duration, a documented ICU stay shorter than 24 h, or a documented death occurring on or before the 24 h landmark were excluded. Physiological measurements recorded during the first 24 h were summarized to construct the predictor set. Five classification algorithms—Logistic Regression, Random Forest, XGBoost, LightGBM, and CatBoost—were developed using a stratified 80:20 train–test split. Performance was evaluated using discrimination, precision–recall, calibration, threshold-dependent classification measures, and decision-curve analysis. Bootstrap resampling was used to estimate 95% confidence intervals and pairwise performance differences. Shapley Additive Explanations (SHAPs) were used to interpret the XGBoost model. Results: The revised cohort included 46,259 survivors and 5339 patients with an in-hospital mortality outcome. XGBoost achieved the numerically highest AUROC of 0.803 (95% CI: 0.790–0.815) and AUPRC of 0.341 (95% CI: 0.313–0.369), with a Brier score of 0.080 (95% CI: 0.076–0.084). Paired bootstrap comparisons showed that XGBoost had a significantly higher AUROC than Random Forest and LightGBM, but not CatBoost. Its AUPRC was significantly higher than that of Random Forest but did not differ significantly from LightGBM or CatBoost, while its Brier score was significantly lower than that of Random Forest but did not differ significantly from LightGBM or CatBoost. At the default probability threshold of 0.50, XGBoost achieved high specificity (0.990) but low recall (0.091), whereas lowering the threshold increased mortality detection at the cost of additional false-positive classifications. Conclusions: Routinely recorded physiological information accumulated during the first 24 h of ICU care supported an interpretable machine learning framework with useful discrimination and probability estimation for subsequent in-hospital mortality. XGBoost achieved the highest numerical performance, although differences among the gradient-boosting models were generally small and not significant for all evaluated metrics. Threshold-sensitivity analysis demonstrated that classification behavior depended strongly on the selected operating threshold. Independent multicenter validation and prospective clinical evaluation are required before routine deployment.
Authors
- In Cheol Jeong (ORCID: https://orcid.org/0000-0001-9314-5601)
- Abdul Karim (ORCID: https://orcid.org/0000-0003-2190-7210)
- Jinwon Kim (ORCID: https://orcid.org/0000-0001-8620-161X)
Institutions
- Hallym University (KR)
- Icahn School of Medicine at Mount Sinai (US)
Publication Details
- Journal
- Diagnostics
- Published
- 2026-09-15
- DOI
- https://doi.org/10.3390/diagnostics16182982
- Primary Topic
- Sepsis Diagnosis and Treatment
- Type
- article
- Field-Weighted Citation Impact
- 0.00