Explainable Machine Learning for Prediction of Future High-Severity States Using Longitudinal APACHE-II Trajectories: Internal Validation and Cross-Cohort Portability Assessment
Background: Longitudinal severity assessment may provide more informative risk stratification than reliance on a single admission score in intensive care. This study developed and internally validated an explainable machine learning framework for predicting a subsequent APACHE-II-defined high-severity state using repeated APACHE-II assessments in intensive care unit (ICU) patients receiving total parenteral nutrition (TPN). The endpoint was defined as an APACHE-II score of ≥ 20 at the final assessment (T10), while predictor information was restricted to measurements obtained at T0–T9. Methods: A retrospective institutional cohort of 844 adult ICU patients was analyzed, of whom 214 (25.4%) met the predefined high-severity endpoint. Four principal feature representations were evaluated: admission APACHE-II, longitudinal APACHE-II information combining the raw T0–T9 sequence with derived trajectory descriptors, longitudinal laboratory information, and multimodal combinations including baseline characteristics. Six machine learning algorithms were compared using stratified five-fold cross-validation and independent hold-out testing. Additional analyses included simpler APACHE-II comparators, sequential observation truncation, nested cross-validation, calibration assessment, threshold sensitivity analysis, and SHAP-based model interpretation. Because equivalent longitudinal APACHE-II measurements and an equivalent endpoint were unavailable in eICU, a separate harmonized XGBoost model was used only for exploratory cross-cohort portability assessment. Results: The longitudinal APACHE-II CatBoost model achieved the highest performance, with a cross-validated ROC-AUC of 0.924 and an independent hold-out ROC-AUC of 0.921 (95% CI: 0.878–0.960). Hold-out calibration was good (Brier score = 0.095, calibration intercept = − 0.03, slope = 0.98), and nested cross-validation yielded a pooled out-of-fold ROC-AUC of 0.908 and a mean outer-fold ROC-AUC of 0.920 ± 0.038. The latest APACHE-II assessment alone (T9; ROC-AUC = 0.553), change from baseline (0.576), ordinal slope (0.539), and logistic regression using engineered APACHE-II descriptors (0.633) performed substantially below the full longitudinal CatBoost model. Sequential truncation showed that discrimination was retained after removing later observations, although performance varied non-monotonically across truncated sequences. SHAP analysis showed that multiple APACHE-II observations contributed prominently to model predictions, with trajectory-derived descriptors providing complementary contributions. Adding laboratory trajectories or baseline characteristics did not improve discrimination over the longitudinal APACHE-II representation. Conclusions: Nonlinear modeling of repeated APACHE-II assessments provided strong internal discrimination of a subsequent APACHE-II-defined high-severity state in this TPN-treated ICU cohort and substantially outperformed single-score and simpler longitudinal comparators. However, the predictors and endpoint share the APACHE-II construct, observation indices were sequential rather than standardized clock-time intervals, and the primary model could not undergo conventional external validation in eICU. These findings therefore support the internal predictive value of longitudinal APACHE-II modeling for this specific severity-state task and warrant prospective multicenter evaluation using standardized timing and independent clinical outcomes.
Authors
- Fatma Çelik (ORCID: https://orcid.org/0000-0003-0192-0151)
- İrem Akpolat (ORCID: https://orcid.org/0000-0002-6136-0021)
Institutions
- Biruni University (TR)
- Okan University (TR)
Publication Details
- Journal
- Diagnostics
- Published
- 2026-09-25
- DOI
- https://doi.org/10.3390/diagnostics16193118
- Primary Topic
- Clinical Nutrition and Gastroenterology
- Type
- article
- Field-Weighted Citation Impact
- 0.00