Interpretable Machine Learning for In-Hospital Mortality Prediction in ICU Patients Using First-24-Hour Routine Vital Signs: A SHAP-Based MIMIC-IV Study

Background: Identification of ICU patients at increased risk of death may support clinical assessment and prioritization of care. However, many mortality prediction approaches incorporate extensive laboratory information, conventional severity scores, or computationally complex models that may limit straightforward interpretation at the bedside. This study investigated whether routinely recorded physiological measurements collected during the first 24 h of ICU care could support an interpretable machine learning approach to subsequent in-hospital mortality prediction. Methods: We performed a retrospective analysis of 51,598 adult first ICU admissions from MIMIC-IV using a strict 24 h landmark design. Patients with a missing ICU duration, a documented ICU stay shorter than 24 h, or a documented death occurring on or before the 24 h landmark were excluded. Physiological measurements recorded during the first 24 h were summarized to construct the predictor set. Five classification algorithms—Logistic Regression, Random Forest, XGBoost, LightGBM, and CatBoost—were developed using a stratified 80:20 train–test split. Performance was evaluated using discrimination, precision–recall, calibration, threshold-dependent classification measures, and decision-curve analysis. Bootstrap resampling was used to estimate 95% confidence intervals and pairwise performance differences. Shapley Additive Explanations (SHAPs) were used to interpret the XGBoost model. Results: The revised cohort included 46,259 survivors and 5339 patients with an in-hospital mortality outcome. XGBoost achieved the numerically highest AUROC of 0.803 (95% CI: 0.790–0.815) and AUPRC of 0.341 (95% CI: 0.313–0.369), with a Brier score of 0.080 (95% CI: 0.076–0.084). Paired bootstrap comparisons showed that XGBoost had a significantly higher AUROC than Random Forest and LightGBM, but not CatBoost. Its AUPRC was significantly higher than that of Random Forest but did not differ significantly from LightGBM or CatBoost, while its Brier score was significantly lower than that of Random Forest but did not differ significantly from LightGBM or CatBoost. At the default probability threshold of 0.50, XGBoost achieved high specificity (0.990) but low recall (0.091), whereas lowering the threshold increased mortality detection at the cost of additional false-positive classifications. Conclusions: Routinely recorded physiological information accumulated during the first 24 h of ICU care supported an interpretable machine learning framework with useful discrimination and probability estimation for subsequent in-hospital mortality. XGBoost achieved the highest numerical performance, although differences among the gradient-boosting models were generally small and not significant for all evaluated metrics. Threshold-sensitivity analysis demonstrated that classification behavior depended strongly on the selected operating threshold. Independent multicenter validation and prospective clinical evaluation are required before routine deployment.

Authors

Institutions

Publication Details

Journal
Diagnostics
Published
2026-09-15
DOI
https://doi.org/10.3390/diagnostics16182982
Primary Topic
Sepsis Diagnosis and Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Interpretable Machine Learning for In-Hospital Mortality Prediction in ICU Patients Using First-24-Hour Routine Vital Signs: A SHAP-Based MIMIC-IV Study

In Cheol Jeong, Abdul Karim, Jinwon Kim
Diagnostics
Sepsis Diagnosis and Treatment
article

Interpretable Machine Learning for In-Hospital Mortality Prediction in ICU Patients Using First-24-Hour Routine Vital Signs: A SHAP-Based MIMIC-IV Study

In Cheol Jeong, Abdul Karim, Jinwon Kim
article en

Abstract

Background: Identification of ICU patients at increased risk of death may support clinical assessment and prioritization of care. However, many mortality prediction approaches incorporate extensive laboratory information, conventional severity scores, or computationally complex models that may limit straightforward interpretation at the bedside. This study investigated whether routinely recorded physiological measurements collected during the first 24 h of ICU care could support an interpretable machine learning approach to subsequent in-hospital mortality prediction. Methods: We performed a retrospective analysis of 51,598 adult first ICU admissions from MIMIC-IV using a strict 24 h landmark design. Patients with a missing ICU duration, a documented ICU stay shorter than 24 h, or a documented death occurring on or before the 24 h landmark were excluded. Physiological measurements recorded during the first 24 h were summarized to construct the predictor set. Five classification algorithms—Logistic Regression, Random Forest, XGBoost, LightGBM, and CatBoost—were developed using a stratified 80:20 train–test split. Performance was evaluated using discrimination, precision–recall, calibration, threshold-dependent classification measures, and decision-curve analysis. Bootstrap resampling was used to estimate 95% confidence intervals and pairwise performance differences. Shapley Additive Explanations (SHAPs) were used to interpret the XGBoost model. Results: The revised cohort included 46,259 survivors and 5339 patients with an in-hospital mortality outcome. XGBoost achieved the numerically highest AUROC of 0.803 (95% CI: 0.790–0.815) and AUPRC of 0.341 (95% CI: 0.313–0.369), with a Brier score of 0.080 (95% CI: 0.076–0.084). Paired bootstrap comparisons showed that XGBoost had a significantly higher AUROC than Random Forest and LightGBM, but not CatBoost. Its AUPRC was significantly higher than that of Random Forest but did not differ significantly from LightGBM or CatBoost, while its Brier score was significantly lower than that of Random Forest but did not differ significantly from LightGBM or CatBoost. At the default probability threshold of 0.50, XGBoost achieved high specificity (0.990) but low recall (0.091), whereas lowering the threshold increased mortality detection at the cost of additional false-positive classifications. Conclusions: Routinely recorded physiological information accumulated during the first 24 h of ICU care supported an interpretable machine learning framework with useful discrimination and probability estimation for subsequent in-hospital mortality. XGBoost achieved the highest numerical performance, although differences among the gradient-boosting models were generally small and not significant for all evaluated metrics. Threshold-sensitivity analysis demonstrated that classification behavior depended strongly on the selected operating threshold. Independent multicenter validation and prospective clinical evaluation are required before routine deployment.

DiagnosticsVol. 16(18)
Hallym University (KR), Icahn School of Medicine at Mount Sinai (US)
Reduced inequalities
Openalex Percentile: Top 10%
Sepsis Diagnosis and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.