Explainable Machine Learning for Prediction of Future High-Severity States Using Longitudinal APACHE-II Trajectories: Internal Validation and Cross-Cohort Portability Assessment

Background: Longitudinal severity assessment may provide more informative risk stratification than reliance on a single admission score in intensive care. This study developed and internally validated an explainable machine learning framework for predicting a subsequent APACHE-II-defined high-severity state using repeated APACHE-II assessments in intensive care unit (ICU) patients receiving total parenteral nutrition (TPN). The endpoint was defined as an APACHE-II score of ≥ 20 at the final assessment (T10), while predictor information was restricted to measurements obtained at T0–T9. Methods: A retrospective institutional cohort of 844 adult ICU patients was analyzed, of whom 214 (25.4%) met the predefined high-severity endpoint. Four principal feature representations were evaluated: admission APACHE-II, longitudinal APACHE-II information combining the raw T0–T9 sequence with derived trajectory descriptors, longitudinal laboratory information, and multimodal combinations including baseline characteristics. Six machine learning algorithms were compared using stratified five-fold cross-validation and independent hold-out testing. Additional analyses included simpler APACHE-II comparators, sequential observation truncation, nested cross-validation, calibration assessment, threshold sensitivity analysis, and SHAP-based model interpretation. Because equivalent longitudinal APACHE-II measurements and an equivalent endpoint were unavailable in eICU, a separate harmonized XGBoost model was used only for exploratory cross-cohort portability assessment. Results: The longitudinal APACHE-II CatBoost model achieved the highest performance, with a cross-validated ROC-AUC of 0.924 and an independent hold-out ROC-AUC of 0.921 (95% CI: 0.878–0.960). Hold-out calibration was good (Brier score = 0.095, calibration intercept = − 0.03, slope = 0.98), and nested cross-validation yielded a pooled out-of-fold ROC-AUC of 0.908 and a mean outer-fold ROC-AUC of 0.920 ± 0.038. The latest APACHE-II assessment alone (T9; ROC-AUC = 0.553), change from baseline (0.576), ordinal slope (0.539), and logistic regression using engineered APACHE-II descriptors (0.633) performed substantially below the full longitudinal CatBoost model. Sequential truncation showed that discrimination was retained after removing later observations, although performance varied non-monotonically across truncated sequences. SHAP analysis showed that multiple APACHE-II observations contributed prominently to model predictions, with trajectory-derived descriptors providing complementary contributions. Adding laboratory trajectories or baseline characteristics did not improve discrimination over the longitudinal APACHE-II representation. Conclusions: Nonlinear modeling of repeated APACHE-II assessments provided strong internal discrimination of a subsequent APACHE-II-defined high-severity state in this TPN-treated ICU cohort and substantially outperformed single-score and simpler longitudinal comparators. However, the predictors and endpoint share the APACHE-II construct, observation indices were sequential rather than standardized clock-time intervals, and the primary model could not undergo conventional external validation in eICU. These findings therefore support the internal predictive value of longitudinal APACHE-II modeling for this specific severity-state task and warrant prospective multicenter evaluation using standardized timing and independent clinical outcomes.

Authors

Institutions

Publication Details

Journal
Diagnostics
Published
2026-09-25
DOI
https://doi.org/10.3390/diagnostics16193118
Primary Topic
Clinical Nutrition and Gastroenterology
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Explainable Machine Learning for Prediction of Future High-Severity States Using Longitudinal APACHE-II Trajectories: Internal Validation and Cross-Cohort Portability Assessment

Fatma Çelik, İrem Akpolat
Diagnostics
Clinical Nutrition and Gastroenterology
article

Explainable Machine Learning for Prediction of Future High-Severity States Using Longitudinal APACHE-II Trajectories: Internal Validation and Cross-Cohort Portability Assessment

Fatma Çelik, İrem Akpolat
article en

Abstract

Background: Longitudinal severity assessment may provide more informative risk stratification than reliance on a single admission score in intensive care. This study developed and internally validated an explainable machine learning framework for predicting a subsequent APACHE-II-defined high-severity state using repeated APACHE-II assessments in intensive care unit (ICU) patients receiving total parenteral nutrition (TPN). The endpoint was defined as an APACHE-II score of ≥ 20 at the final assessment (T10), while predictor information was restricted to measurements obtained at T0–T9. Methods: A retrospective institutional cohort of 844 adult ICU patients was analyzed, of whom 214 (25.4%) met the predefined high-severity endpoint. Four principal feature representations were evaluated: admission APACHE-II, longitudinal APACHE-II information combining the raw T0–T9 sequence with derived trajectory descriptors, longitudinal laboratory information, and multimodal combinations including baseline characteristics. Six machine learning algorithms were compared using stratified five-fold cross-validation and independent hold-out testing. Additional analyses included simpler APACHE-II comparators, sequential observation truncation, nested cross-validation, calibration assessment, threshold sensitivity analysis, and SHAP-based model interpretation. Because equivalent longitudinal APACHE-II measurements and an equivalent endpoint were unavailable in eICU, a separate harmonized XGBoost model was used only for exploratory cross-cohort portability assessment. Results: The longitudinal APACHE-II CatBoost model achieved the highest performance, with a cross-validated ROC-AUC of 0.924 and an independent hold-out ROC-AUC of 0.921 (95% CI: 0.878–0.960). Hold-out calibration was good (Brier score = 0.095, calibration intercept = − 0.03, slope = 0.98), and nested cross-validation yielded a pooled out-of-fold ROC-AUC of 0.908 and a mean outer-fold ROC-AUC of 0.920 ± 0.038. The latest APACHE-II assessment alone (T9; ROC-AUC = 0.553), change from baseline (0.576), ordinal slope (0.539), and logistic regression using engineered APACHE-II descriptors (0.633) performed substantially below the full longitudinal CatBoost model. Sequential truncation showed that discrimination was retained after removing later observations, although performance varied non-monotonically across truncated sequences. SHAP analysis showed that multiple APACHE-II observations contributed prominently to model predictions, with trajectory-derived descriptors providing complementary contributions. Adding laboratory trajectories or baseline characteristics did not improve discrimination over the longitudinal APACHE-II representation. Conclusions: Nonlinear modeling of repeated APACHE-II assessments provided strong internal discrimination of a subsequent APACHE-II-defined high-severity state in this TPN-treated ICU cohort and substantially outperformed single-score and simpler longitudinal comparators. However, the predictors and endpoint share the APACHE-II construct, observation indices were sequential rather than standardized clock-time intervals, and the primary model could not undergo conventional external validation in eICU. These findings therefore support the internal predictive value of longitudinal APACHE-II modeling for this specific severity-state task and warrant prospective multicenter evaluation using standardized timing and independent clinical outcomes.

DiagnosticsVol. 16(19)
Biruni University (TR), Okan University (TR)
Zero hunger
Openalex Percentile: Top 13%
Clinical Nutrition and Gastroenterology
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.