Development and internal validation of an explainable machine learning model for postoperative pulmonary infection risk stratification in kidney transplant recipients
Kidney transplant recipients are at substantial risk of postoperative pulmonary infection, yet effective tools for postoperative clinical risk stratification remain limited. This study aimed to develop and internally validate an interpretable machine learning model for postoperative pulmonary infection after kidney transplantation. We retrospectively collected clinical data from 410 patients who underwent kidney transplantation at the Organ Transplantation Center of the First Affiliated Hospital of Kunming Medical University. The dataset was randomly divided into training and test sets at a ratio of 7:3. Preoperative and postoperative day 7 clinical and laboratory variables were used to develop a model for clinical risk stratification of postoperative pulmonary infection documented during the index hospitalization. Eight machine learning algorithms, including support vector machine (SVM), gradient boosting machine (GBM), neural network (NNET), extreme gradient boosting (XGBoost), k-nearest neighbors (KNN), adaptive boosting (AdaBoost), light gradient boosting machine (LightGBM), and random forest (RF), were used to build postoperative pulmonary infection risk prediction models. Model performance was evaluated using AUC, accuracy, sensitivity, specificity, precision, and F1 score. Decision curve analysis (DCA) was employed to assess potential clinical utility. Among the eight algorithms, the random forest model showed the best overall classification performance. The area under the curve was 0.866 (95% CI, 0.827–0.906) in the training set and 0.780 (95% CI, 0.675–0.886) in the test set. Decision curve analysis suggested potential clinical utility. Boruta feature selection identified six variables for model development: preoperative white blood cell count (WBC), and postoperative day 7 indicators including serum creatinine (SCr), uric acid (UA), phosphorus (P), neutrophil-to-lymphocyte ratio (NLR), and glucose-to-lymphocyte ratio (GLR). SHAP analysis was subsequently used to interpret the contribution of these variables to the random forest model output. A web-based platform was constructed based on the RF model to facilitate clinical risk assessment. We developed and internally validated an explainable machine learning model for clinical risk stratification of postoperative pulmonary infection in kidney transplant recipients. The model may provide adjunctive information for individualized postoperative assessment, but external validation is required before broader clinical application.
Authors
- Yingjia He (ORCID: https://orcid.org/0000-0002-9071-8926)
- P. Liu (ORCID: https://orcid.org/0009-0003-1751-2184)
- Tao Liu (ORCID: https://orcid.org/0000-0002-4340-5767)
- Fan WenXing
- Zhong Zeng
- HanFei Huang
Institutions
- Kunming Medical University (CN)
- First Affiliated Hospital of Kunming Medical University (CN)
Publication Details
- Journal
- BMC Nephrology
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1186/s12882-026-05352-8
- Primary Topic
- Transplantation: Methods and Outcomes
- Type
- article
- Field-Weighted Citation Impact
- 0.00