Early prediction of sepsis in critically ill trauma patients using machine learning: a MIMIC-IV-based cohort study

Abstract This study aimed to develop and validate machine learning (ML) models for the early prediction of incident sepsis in critically ill trauma patients using the Medical Information Mart for Intensive Care IV (MIMIC-IV) database. Adult trauma patients with an intensive care unit (ICU) stay of at least 24 h were analysed, excluding those with sepsis or suspected infection during the first 24 h. Clinical data from the first 24 h of ICU admission were used at hour 24 to predict incident Sepsis-3 during hours 24–72, and treatment variables and Sequential Organ Failure Assessment (SOFA) scores were excluded to avoid target leakage. Consensus feature selection performed within each training fold retained 19 predictors. Nine ML configurations across seven algorithm families were trained without synthetic oversampling and assessed on a held-out internal evaluation set with 1,000 bootstrap 95% confidence intervals (CIs) and SHapley Additive exPlanations (SHAP). Among 4,043 trauma ICU patients, 450 (11.13%) developed incident sepsis. Discrimination was moderate and similar across models, with the highest area under the receiver operating characteristic curve (AUROC) obtained by polynomial support vector machine (SVM) (0.734, 95% CI 0.680 to 0.788) and linear SVM (0.733, 95% CI 0.678 to 0.785), both superior to SOFA (0.655) and Simplified Acute Physiology Score II (0.668) (DeLong P < 0.01). Sensitivity ranged from 0.589 to 0.689 and specificity from 0.650 to 0.707, gradient boosting showed the best calibration (slope 0.946, intercept -0.073, Brier score 0.092), and decision curve analysis showed positive net benefit at threshold probabilities of 2% to 35%. SHAP analysis identified maximum glucose, mean oxygen saturation, maximum temperature, mean respiratory rate, minimum hemoglobin, and minimum platelet count as the most influential predictors of incident sepsis. ML models based strictly on early physiological and laboratory data provided moderate but well calibrated discrimination for incident sepsis in critically ill trauma patients and outperformed general ICU severity scores. These interpretable models may support risk-based infection surveillance rather than autonomous clinical decision-making, and external validation is warranted before clinical implementation.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-10-08
DOI
https://doi.org/10.1038/s41598-026-73622-4
Primary Topic
Sepsis Diagnosis and Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Early prediction of sepsis in critically ill trauma patients using machine learning: a MIMIC-IV-based cohort study

Guandong Huang, Yanping Yang, Jian Sun, Weixi Zhong et al.
Scientific Reports
Sepsis Diagnosis and Treatment
article

Early prediction of sepsis in critically ill trauma patients using machine learning: a MIMIC-IV-based cohort study

Guandong Huang, Yanping Yang, Jian Sun, Weixi Zhong, Qing Xu
article en

Abstract

Abstract This study aimed to develop and validate machine learning (ML) models for the early prediction of incident sepsis in critically ill trauma patients using the Medical Information Mart for Intensive Care IV (MIMIC-IV) database. Adult trauma patients with an intensive care unit (ICU) stay of at least 24 h were analysed, excluding those with sepsis or suspected infection during the first 24 h. Clinical data from the first 24 h of ICU admission were used at hour 24 to predict incident Sepsis-3 during hours 24–72, and treatment variables and Sequential Organ Failure Assessment (SOFA) scores were excluded to avoid target leakage. Consensus feature selection performed within each training fold retained 19 predictors. Nine ML configurations across seven algorithm families were trained without synthetic oversampling and assessed on a held-out internal evaluation set with 1,000 bootstrap 95% confidence intervals (CIs) and SHapley Additive exPlanations (SHAP). Among 4,043 trauma ICU patients, 450 (11.13%) developed incident sepsis. Discrimination was moderate and similar across models, with the highest area under the receiver operating characteristic curve (AUROC) obtained by polynomial support vector machine (SVM) (0.734, 95% CI 0.680 to 0.788) and linear SVM (0.733, 95% CI 0.678 to 0.785), both superior to SOFA (0.655) and Simplified Acute Physiology Score II (0.668) (DeLong P < 0.01). Sensitivity ranged from 0.589 to 0.689 and specificity from 0.650 to 0.707, gradient boosting showed the best calibration (slope 0.946, intercept -0.073, Brier score 0.092), and decision curve analysis showed positive net benefit at threshold probabilities of 2% to 35%. SHAP analysis identified maximum glucose, mean oxygen saturation, maximum temperature, mean respiratory rate, minimum hemoglobin, and minimum platelet count as the most influential predictors of incident sepsis. ML models based strictly on early physiological and laboratory data provided moderate but well calibrated discrimination for incident sepsis in critically ill trauma patients and outperformed general ICU severity scores. These interpretable models may support risk-based infection surveillance rather than autonomous clinical decision-making, and external validation is warranted before clinical implementation.

Scientific Reports
Openalex Percentile: Top 12%
Sepsis Diagnosis and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.