Development and external validation of a machine learning model for predicting stroke-associated pneumonia in acute ischemic stroke patients receiving pantoprazole

Abstract This study aimed to develop and externally validate a machine learning (ML)-based risk prediction model for stroke-associated pneumonia (SAP) in acute ischemic stroke (AIS) patients receiving pantoprazole. Data from AIS patients receiving pantoprazole were retrospectively extracted from the MIMIC-IV 3.0 database. This retrospective cohort study enrolled 651 AIS patients treated with pantoprazole from the MIMIC-IV 3.0 database, randomly partitioned into a training cohort ( n = 455) and an internal validation cohort ( n = 196) at a 7:3 ratio. The primary outcome was incident SAP. A two-stage feature selection strategy was adopted: LASSO regression for preliminary filtering, followed by bidirectional stepwise multivariable logistic regression. Subsequently, nine ML algorithms—decision tree (DT), k-nearest neighbor (KNN), random forest (RF), support vector machine (SVM), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), categorical boosting (CatBoost), naive Bayes (NB), neural network (NN)—were trained and evaluated via 10-fold cross-validation. Model performance was comprehensively assessed by the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, accuracy, F1 score, Brier score, calibration curve, and decision curve analysis (DCA). Independent external validation was conducted using 580 patients from the eICU-CRD database. The optimal model was selected and interpreted via SHapley Additive exPlanations (SHAP) analysis to elucidate feature importance and decision rationale. Finally, a nomogram was constructed to increase the readability of the predictive results. Age, male sex, platelet count, respiratory rate, and Sepsis-3 status were identified as independent predictors of SAP. The LightGBM model outperformed the other ML algorithms, with an internal validation AUC of 0.671 (95% CI: 0.593–0.749), a sensitivity of 0.544, a specificity of 0.735, an accuracy of 0.640, and an F1 score of 0.562. In the external validation set, the AUC was 0.679 (95% CI: 0.612–0.746), accompanied by acceptable calibration and measurable clinical net benefit within practical clinical risk ranges. SHAP analysis further revealed that Sepsis-3, platelet count, and age were the primary determinants influencing model predictions. The LightGBM model demonstrates moderate discriminative performance in predicting SAP among pantoprazole-treated AIS patients. It may serve as an adjunctive tool for early risk stratification but is insufficient for standalone SAP diagnosis.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-09-28
DOI
https://doi.org/10.1038/s41598-026-73858-0
Primary Topic
Pneumonia and Respiratory Infections
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Development and external validation of a machine learning model for predicting stroke-associated pneumonia in acute ischemic stroke patients receiving pantoprazole

Sufang Yang, Jin Min, Guohua Liu, Lin Chen et al.
Scientific Reports
Pneumonia and Respiratory Infections
article

Development and external validation of a machine learning model for predicting stroke-associated pneumonia in acute ischemic stroke patients receiving pantoprazole

Sufang Yang, Jin Min, Guohua Liu, Lin Chen, Lu Li, Ji Li
article en

Abstract

Abstract This study aimed to develop and externally validate a machine learning (ML)-based risk prediction model for stroke-associated pneumonia (SAP) in acute ischemic stroke (AIS) patients receiving pantoprazole. Data from AIS patients receiving pantoprazole were retrospectively extracted from the MIMIC-IV 3.0 database. This retrospective cohort study enrolled 651 AIS patients treated with pantoprazole from the MIMIC-IV 3.0 database, randomly partitioned into a training cohort ( n = 455) and an internal validation cohort ( n = 196) at a 7:3 ratio. The primary outcome was incident SAP. A two-stage feature selection strategy was adopted: LASSO regression for preliminary filtering, followed by bidirectional stepwise multivariable logistic regression. Subsequently, nine ML algorithms—decision tree (DT), k-nearest neighbor (KNN), random forest (RF), support vector machine (SVM), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), categorical boosting (CatBoost), naive Bayes (NB), neural network (NN)—were trained and evaluated via 10-fold cross-validation. Model performance was comprehensively assessed by the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, accuracy, F1 score, Brier score, calibration curve, and decision curve analysis (DCA). Independent external validation was conducted using 580 patients from the eICU-CRD database. The optimal model was selected and interpreted via SHapley Additive exPlanations (SHAP) analysis to elucidate feature importance and decision rationale. Finally, a nomogram was constructed to increase the readability of the predictive results. Age, male sex, platelet count, respiratory rate, and Sepsis-3 status were identified as independent predictors of SAP. The LightGBM model outperformed the other ML algorithms, with an internal validation AUC of 0.671 (95% CI: 0.593–0.749), a sensitivity of 0.544, a specificity of 0.735, an accuracy of 0.640, and an F1 score of 0.562. In the external validation set, the AUC was 0.679 (95% CI: 0.612–0.746), accompanied by acceptable calibration and measurable clinical net benefit within practical clinical risk ranges. SHAP analysis further revealed that Sepsis-3, platelet count, and age were the primary determinants influencing model predictions. The LightGBM model demonstrates moderate discriminative performance in predicting SAP among pantoprazole-treated AIS patients. It may serve as an adjunctive tool for early risk stratification but is insufficient for standalone SAP diagnosis.

Scientific Reports
Peace, Justice and strong institutions
Openalex Percentile: Top 11%
Pneumonia and Respiratory Infections
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.