Development and a single-centre temporal validation of a machine learning-based model for early prediction of acute pancreatitis-associated liver injury

Acute pancreatitis (AP) is a common acute abdominal inflammatory disorder characterized by premature activation of pancreatic enzymes, leading to pancreatic autodigestion, edema, hemorrhage, and even necrosis. The liver, as the first portal organ receiving pancreatic venous drainage, is one of the most frequently affected extra-pancreatic organs in AP. AP-associated liver injury (LI) is closely correlated with the intensity of systemic inflammation, the development of multiple organ dysfunction, and increased mortality. While machine learning (ML) models have been successfully applied to predict various AP-related complications such as acute respiratory distress syndrome, acute kidney injury, and infected pancreatic necrosis, effective tools for early prediction of LI before its clinical onset remain lacking. We retrospectively collected data from 1,184 patients diagnosed with AP at the First Affiliated Hospital of Henan University of Science and Technology between January 2018 and December 2023. After applying strict inclusion and exclusion criteria, 1,124 patients were included. The development cohort comprised 929 patients admitted between January 2018 and December 2022, randomly split into a training set (n = 650) and an internal test set (n = 279) in a 7:3 ratio. An independent temporal validation set consisted of 195 patients admitted from January to December 2023. Forty-nine clinical and laboratory parameters were collected within 24 h of admission to predict LI occurring after the first 24 h of hospitalization. Patients with any ALT/AST elevation > 3×ULN within the first 24 h were not counted as LI events; the outcome was strictly defined as incident LI emerging after 24 h. In the training set, LASSO regression was used to identify key predictive factors, followed by the construction of nine machine learning models: logistic regression (LR), decision trees (DT), random forests (RF), XGBoost, and LightGBM, among others. All feature selection and hyperparameter tuning were nested within 5-fold cross-validation on the training set only, with no use of test or validation set information. Model performance was primarily assessed using the area under the receiver operating characteristic curve (AUC), calibration metrics (Brier score, calibration intercept/slope), and decision-curve analysis (DCA). The SHAP method was employed for interpretability analysis of the optimal model. LASSO regression identified ten key predictors: amylase (AMY), lipase (LIP), total cholesterol (CHO), albumin (ALB), neutrophil percentage (N/W), serum calcium (Ca), alanine aminotransferase (ALT), gender, total bilirubin (TBIL), and drinking history. Among the nine models, the random forest (RF) model demonstrated superior predictive performance and temporal generalizability, achieving an AUC of 0.972 (95% CI: 0.954–0.989) with an accuracy of 0.921 in the internal test set, and an AUC of 0.935 (95% CI: 0.902–0.969) with an accuracy of 0.862 in the temporal validation set. The RF model showed a training–test AUC gap of 0.028 (1.000 vs. 0.972), indicating modest overfitting. Decision-curve analysis confirmed net clinical benefit across a range of thresholds. SHAP analysis revealed that pancreatic enzymes (AMY, LIP) and markers reflecting hepatic synthetic function (ALB, CHO) were the primary predictors influencing the model’s output. Calibration was acceptable, with Brier scores of 0.070 and 0.101 in the test and validation sets, respectively; calibration slopes were 2.153 (test) and 1.367 (validation), indicating over-dispersion of predicted probabilities (over-confidence) that would benefit from post-hoc recalibration before clinical use. An RF model built using routine clinical data collected within 24 h of admission can effectively predict the risk of liver injury in AP patients and demonstrates promising temporal generalizability. The model offers a potentially useful tool for early identification of high-risk patients; however, prospective external validation and probability recalibration are required before clinical implementation.

Authors

Institutions

Publication Details

Journal
BMC Gastroenterology
Published
2026-09-26
DOI
https://doi.org/10.1186/s12876-026-05204-7
Primary Topic
Pancreatitis Pathology and Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Development and a single-centre temporal validation of a machine learning-based model for early prediction of acute pancreatitis-associated liver injury

Hailong Feng, Keyang Wang, Cunbo Yu, ZHANG Yingjian et al.
BMC Gastroenterology
Pancreatitis Pathology and Treatment
article

Development and a single-centre temporal validation of a machine learning-based model for early prediction of acute pancreatitis-associated liver injury

Hailong Feng, Keyang Wang, Cunbo Yu, ZHANG Yingjian, Xueli Yao, Ping Wang
article en

Abstract

Acute pancreatitis (AP) is a common acute abdominal inflammatory disorder characterized by premature activation of pancreatic enzymes, leading to pancreatic autodigestion, edema, hemorrhage, and even necrosis. The liver, as the first portal organ receiving pancreatic venous drainage, is one of the most frequently affected extra-pancreatic organs in AP. AP-associated liver injury (LI) is closely correlated with the intensity of systemic inflammation, the development of multiple organ dysfunction, and increased mortality. While machine learning (ML) models have been successfully applied to predict various AP-related complications such as acute respiratory distress syndrome, acute kidney injury, and infected pancreatic necrosis, effective tools for early prediction of LI before its clinical onset remain lacking. We retrospectively collected data from 1,184 patients diagnosed with AP at the First Affiliated Hospital of Henan University of Science and Technology between January 2018 and December 2023. After applying strict inclusion and exclusion criteria, 1,124 patients were included. The development cohort comprised 929 patients admitted between January 2018 and December 2022, randomly split into a training set (n = 650) and an internal test set (n = 279) in a 7:3 ratio. An independent temporal validation set consisted of 195 patients admitted from January to December 2023. Forty-nine clinical and laboratory parameters were collected within 24 h of admission to predict LI occurring after the first 24 h of hospitalization. Patients with any ALT/AST elevation > 3×ULN within the first 24 h were not counted as LI events; the outcome was strictly defined as incident LI emerging after 24 h. In the training set, LASSO regression was used to identify key predictive factors, followed by the construction of nine machine learning models: logistic regression (LR), decision trees (DT), random forests (RF), XGBoost, and LightGBM, among others. All feature selection and hyperparameter tuning were nested within 5-fold cross-validation on the training set only, with no use of test or validation set information. Model performance was primarily assessed using the area under the receiver operating characteristic curve (AUC), calibration metrics (Brier score, calibration intercept/slope), and decision-curve analysis (DCA). The SHAP method was employed for interpretability analysis of the optimal model. LASSO regression identified ten key predictors: amylase (AMY), lipase (LIP), total cholesterol (CHO), albumin (ALB), neutrophil percentage (N/W), serum calcium (Ca), alanine aminotransferase (ALT), gender, total bilirubin (TBIL), and drinking history. Among the nine models, the random forest (RF) model demonstrated superior predictive performance and temporal generalizability, achieving an AUC of 0.972 (95% CI: 0.954–0.989) with an accuracy of 0.921 in the internal test set, and an AUC of 0.935 (95% CI: 0.902–0.969) with an accuracy of 0.862 in the temporal validation set. The RF model showed a training–test AUC gap of 0.028 (1.000 vs. 0.972), indicating modest overfitting. Decision-curve analysis confirmed net clinical benefit across a range of thresholds. SHAP analysis revealed that pancreatic enzymes (AMY, LIP) and markers reflecting hepatic synthetic function (ALB, CHO) were the primary predictors influencing the model’s output. Calibration was acceptable, with Brier scores of 0.070 and 0.101 in the test and validation sets, respectively; calibration slopes were 2.153 (test) and 1.367 (validation), indicating over-dispersion of predicted probabilities (over-confidence) that would benefit from post-hoc recalibration before clinical use. An RF model built using routine clinical data collected within 24 h of admission can effectively predict the risk of liver injury in AP patients and demonstrates promising temporal generalizability. The model offers a potentially useful tool for early identification of high-risk patients; however, prospective external validation and probability recalibration are required before clinical implementation.

BMC Gastroenterology
Henan University of Science and Technology (CN), First Affiliated Hospital of Henan University of Science and Technology (CN)
Good health and well-being
Openalex Percentile: Top 8%
Pancreatitis Pathology and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.