Machine Learning–Based Prediction of Culture-Confirmed Neonatal Sepsis in a Tertiary Neonatal Intensive Care Unit: Retrospective Cohort Study
Abstract Background Neonatal sepsis remains a major cause of neonatal morbidity and mortality in low- and middle-income countries (LMICs). Early diagnosis is challenging because of nonspecific clinical manifestations and delays in laboratory confirmation. Machine learning (ML) approaches using structured electronic health record (EHR) data may improve early risk stratification in neonatal intensive care units (NICUs). Objective This study aimed to evaluate ML models for predicting culture-confirmed neonatal sepsis among neonates admitted to a tertiary NICU in Jordan, with the objective of addressing diagnostic gaps in resource-limited settings. Specifically, we aimed to identify key predictors through feature importance analysis, evaluate model performance with class-imbalanced data, and propose strategies to improve interpretability and generalizability in LMICs. Methods A retrospective cohort study was conducted using structured EHRs of 3274 neonates admitted to a tertiary NICU in Jordan between 2018 and 2024. Neonates who underwent blood culture testing were included. The dataset was divided into training (n=2619, 80%) and testing (n=655, 20%) subsets using stratified sampling. Three ML models—Extreme Gradient Boosting (XGBoost), decision trees, and neural networks—were trained using clinical, laboratory, and demographic variables. Class imbalance was addressed using the synthetic minority oversampling technique (SMOTE) applied to the training dataset. Model performance was evaluated using accuracy, sensitivity, specificity, and the area under the receiver operating characteristic curve (AUC). Results Among 3274 neonates included in the study, the XGBoost model demonstrated the best predictive performance on the independent test set (n=655, 20%), achieving an accuracy of 94% (616/655 correct predictions, 95% CI 92% to 96%), sensitivity of 98% (95/97 sepsis cases correctly identified, 95% CI 96% to 99%), and an AUC of 0.98 (95% CI 0.97 to 0.99). Decision trees provided interpretable classification rules with moderate performance, whereas neural networks showed lower discriminative ability, with an AUC of 0.81 (95% CI 0.78 to 0.84). Important predictive features included C-reactive protein, platelet count, and gestational age. Conclusions XGBoost demonstrated strong predictive performance in this retrospective cohort, supporting its potential as a foundation for future prospective clinical decision support tools. External validation and prospective studies are required before clinical implementation.
Authors
- Shahd H Rihan (ORCID: https://orcid.org/0000-0003-1235-3228)
- Eman Badran (ORCID: https://orcid.org/0000-0003-1124-592X)
- Abdulrahman E. Alhanbali (ORCID: https://orcid.org/0000-0003-1860-4294)
- Arwa Al Anber (ORCID: https://orcid.org/0000-0001-8460-0748)
- Alaa T. Al Ghazo (ORCID: https://orcid.org/0000-0001-7029-4487)
- Hala Al-Jaberi (ORCID: https://orcid.org/0000-0002-3337-9657)
- Areej Sharaqa (ORCID: https://orcid.org/0009-0001-9118-2342)
- Loiy T. Algazo (ORCID: https://orcid.org/0009-0002-6526-1757)
- Oraib Al-Smadi (ORCID: https://orcid.org/0009-0001-6823-7392)
- Taimein Yacoub (ORCID: https://orcid.org/0009-0009-0808-3086)
- Shatha Al Jaberi (ORCID: https://orcid.org/0009-0004-0867-1316)
- Lena Abu-Argoub (ORCID: https://orcid.org/0000-0002-2386-6443)
Publication Details
- Journal
- JMIR Medical Informatics
- Published
- 2026-09-14
- DOI
- https://doi.org/10.2196/88732
- Primary Topic
- Neonatal and Maternal Infections
- Type
- article
- Field-Weighted Citation Impact
- 0.00