Development and interpretability of a machine learning-derived model for predicting the risk of subsequent short-term infection in patients undergoing maintenance hemodialysis: a retrospective cohort study

Abstract Background Infections are a major cause of mortality in hemodialysis patients, but early risk stratification remains difficult. This study developed and validated an interpretable ML model using routine baseline laboratory data to predict short-term infection risk. Methods A retrospective cohort of 622 hemodialysis patients with uremia was analyzed. The primary outcome was a confirmed infectious event (infection, sepsis, or pneumonia) occurring within a 2-month follow-up. To rigorously prevent data leakage, the cohort was strictly partitioned into training (70%) and test (30%) sets prior to missing value imputation and any downstream preprocessing. A LightGBM-based forward feature selection was executed exclusively within the training set. Nine ML algorithms were evaluated, and model robustness was validated via nested cross-validation, repeated cross-validation, and bootstrap optimism correction. Interpretability was achieved using SHapley Additive exPlanations (SHAP). Results Subsequent infections occurred in 157 patients (25.2%). A parsimonious six-biomarker signature (CysC, HCT, TP, MONO_pct, MYO, ADA) was identified. The K-Nearest Neighbors (KNN) and LightGBM models demonstrated robust discriminative performance (Test AUROC: 0.727 and 0.719, respectively). Nested cross-validation confirmed generalization stability (e.g., ANN mean AUC: 0.722). SHAP analysis revealed elevated Cystatin C and lowered Hematocrit as the strongest contributors to the model’s risk ranking. Although the model showed high ranking ability, the default 0.5 threshold yielded low sensitivity; a prespecified F2‑optimized threshold of 0.30 improved sensitivity to 0.723 in the test set. Conclusion This interpretable machine learning‑based predictive model, using routinely available laboratory data, may serve as a potential adjunctive decision‑support tool for stratifying short‑term infection risk. External validation is required before clinical implementation.

Authors

Publication Details

Journal
BMC Nephrology
Published
2026-10-08
DOI
https://doi.org/10.1186/s12882-026-05419-6
Primary Topic
Dialysis and Renal Disease Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Development and interpretability of a machine learning-derived model for predicting the risk of subsequent short-term infection in patients undergoing maintenance hemodialysis: a retrospective cohort study

苏恒学, Feng Li, Jie Huang, Xuefei Liang et al.
BMC Nephrology
Dialysis and Renal Disease Management
article

Development and interpretability of a machine learning-derived model for predicting the risk of subsequent short-term infection in patients undergoing maintenance hemodialysis: a retrospective cohort study

苏恒学, Feng Li, Jie Huang, Xuefei Liang, Yongzhao Li, Ronglan Yang, Xianhong Huang, Liqing Li
article en

Abstract

Abstract Background Infections are a major cause of mortality in hemodialysis patients, but early risk stratification remains difficult. This study developed and validated an interpretable ML model using routine baseline laboratory data to predict short-term infection risk. Methods A retrospective cohort of 622 hemodialysis patients with uremia was analyzed. The primary outcome was a confirmed infectious event (infection, sepsis, or pneumonia) occurring within a 2-month follow-up. To rigorously prevent data leakage, the cohort was strictly partitioned into training (70%) and test (30%) sets prior to missing value imputation and any downstream preprocessing. A LightGBM-based forward feature selection was executed exclusively within the training set. Nine ML algorithms were evaluated, and model robustness was validated via nested cross-validation, repeated cross-validation, and bootstrap optimism correction. Interpretability was achieved using SHapley Additive exPlanations (SHAP). Results Subsequent infections occurred in 157 patients (25.2%). A parsimonious six-biomarker signature (CysC, HCT, TP, MONO_pct, MYO, ADA) was identified. The K-Nearest Neighbors (KNN) and LightGBM models demonstrated robust discriminative performance (Test AUROC: 0.727 and 0.719, respectively). Nested cross-validation confirmed generalization stability (e.g., ANN mean AUC: 0.722). SHAP analysis revealed elevated Cystatin C and lowered Hematocrit as the strongest contributors to the model’s risk ranking. Although the model showed high ranking ability, the default 0.5 threshold yielded low sensitivity; a prespecified F2‑optimized threshold of 0.30 improved sensitivity to 0.723 in the test set. Conclusion This interpretable machine learning‑based predictive model, using routinely available laboratory data, may serve as a potential adjunctive decision‑support tool for stratifying short‑term infection risk. External validation is required before clinical implementation.

BMC Nephrology
Openalex Percentile: Top 12%
Dialysis and Renal Disease Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Development and interpretability of a machine learning-derived model for predicting the risk of subsequent short-term infection in patients undergoing maintenance hemodialysis: a retrospective cohort study — 苏恒学, Feng Li, et al. · BMC Nephrology (2026) | TGRS Research Map | TGRS