Development and external validation of a machine learning model for predicting stroke-associated pneumonia in acute ischemic stroke patients receiving pantoprazole
Abstract This study aimed to develop and externally validate a machine learning (ML)-based risk prediction model for stroke-associated pneumonia (SAP) in acute ischemic stroke (AIS) patients receiving pantoprazole. Data from AIS patients receiving pantoprazole were retrospectively extracted from the MIMIC-IV 3.0 database. This retrospective cohort study enrolled 651 AIS patients treated with pantoprazole from the MIMIC-IV 3.0 database, randomly partitioned into a training cohort ( n = 455) and an internal validation cohort ( n = 196) at a 7:3 ratio. The primary outcome was incident SAP. A two-stage feature selection strategy was adopted: LASSO regression for preliminary filtering, followed by bidirectional stepwise multivariable logistic regression. Subsequently, nine ML algorithms—decision tree (DT), k-nearest neighbor (KNN), random forest (RF), support vector machine (SVM), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), categorical boosting (CatBoost), naive Bayes (NB), neural network (NN)—were trained and evaluated via 10-fold cross-validation. Model performance was comprehensively assessed by the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, accuracy, F1 score, Brier score, calibration curve, and decision curve analysis (DCA). Independent external validation was conducted using 580 patients from the eICU-CRD database. The optimal model was selected and interpreted via SHapley Additive exPlanations (SHAP) analysis to elucidate feature importance and decision rationale. Finally, a nomogram was constructed to increase the readability of the predictive results. Age, male sex, platelet count, respiratory rate, and Sepsis-3 status were identified as independent predictors of SAP. The LightGBM model outperformed the other ML algorithms, with an internal validation AUC of 0.671 (95% CI: 0.593–0.749), a sensitivity of 0.544, a specificity of 0.735, an accuracy of 0.640, and an F1 score of 0.562. In the external validation set, the AUC was 0.679 (95% CI: 0.612–0.746), accompanied by acceptable calibration and measurable clinical net benefit within practical clinical risk ranges. SHAP analysis further revealed that Sepsis-3, platelet count, and age were the primary determinants influencing model predictions. The LightGBM model demonstrates moderate discriminative performance in predicting SAP among pantoprazole-treated AIS patients. It may serve as an adjunctive tool for early risk stratification but is insufficient for standalone SAP diagnosis.
Authors
- Sufang Yang
- Jin Min (ORCID: https://orcid.org/0009-0000-8747-0731)
- Guohua Liu
- Lin Chen
- Lu Li
- Ji Li
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1038/s41598-026-73858-0
- Primary Topic
- Pneumonia and Respiratory Infections
- Type
- article
- Field-Weighted Citation Impact
- 0.00