Development and validation of an explainable machine learning model for predicting multiple organ failure in patients with acute pancreatitis: a multicenter cohort study

Acute pancreatitis can lead to a serious and life-threatening situation known as multiple organ failure (MOF). Timely and precise forecasting and detection of MOF are essential. Predictive models utilizing machine learning have shown potential in forecasting MOF in emergency situations. Nevertheless, their use for predicting MOF specifically in patients with acute pancreatitis is still not widespread. This research seeks to create and confirm a machine learning-based model for predicting MOF in individuals suffering from acute pancreatitis. This research utilized two retrospective cohorts for the purposes of developing, validating, and testing the model. The derivation and validation cohorts were sourced from the MIMIC-IV database, which was divided randomly into two segments (70% for model development and 30% for internal validation). For external validation, a retrospective cohort from The Department of Hepatobiliary and Pancreatic Surgery at the Third Hospital of Shanxi Medical University (DHPS cohort) was used. After performing several data preprocessing techniques, including interpolation and standardization, we applied four methods for feature selection, ultimately identifying 15 key features. Seven machine learning algorithms were utilized to create predictive models, with their performance assessed through various metrics such as ROC (receiver operating characteristic), DCA (decision curve analysis), PRC (precision recall curve), calibration curve, and confusion matrice. The final model's interpretation was carried out using SHAP technology, and a web-based risk calculator was created for use in clinical settings. Finally, the final model was used to compare with the SOFA score and APACHE Ⅱ score, respectively, to determine its clinical value. The model was created utilizing data from the MIMIC-Ⅳ database (n=582) and underwent validation and testing with both the MIMIC-Ⅳ (n=250) and The Department of Hepatobiliary and Pancreatic Surgery at the Third Hospital of Shanxi Medical University (n=474) datasets. We identified the best feature combination through four different selection techniques and seven machine learning algorithms. A variable that was recognized by all four selection methods was included in the model development process. The most effective model, CatBoost, was built using 15 easily accessible admission features, including INR (international normalized ratio), Cr (creatinine), PO2 (partial pressure of oxygen), Nbps (non-invasive blood pressure systolic), BIL (bilirubin), HTN (hypertension), PH (potential of hydrogen), LAC (lactate), Hb (hemoglobin), CKD (chronic kidney disease), T (temperature), ALT (alanine aminotransferase), HR (heart rate), Na + (sodion), ABX (antibiotic). During validation, it achieved an AUC value of 0.939, an F1 score of 0.795, and an accuracy rate of 0.856. Subsequently, hyperparameter optimization was conducted on the derivation cohort via 5-fold cross-validation and grid search. The final model was assessed on the test dataset, yielding an AUC of 0.855, an F1 score of 0.658, and an accuracy of 0.761. Additionally, SHAP analysis indicated that INR, Cr, and PO2 are the three most significant variables affecting the model's predictions. This model has been transformed into an online clinical tool to enhance its application in healthcare environments. In validation, CatBoost demonstrated superior discrimination (AUC: 0.941, 95%CI: 0.914–0.968) compared to both APACHE II (0.802, 95%CI: 0.747–0.858) and SOFA (0.819, 95%CI: 0.765–0.874) through ROC and decision curve analyses. The CatBoost model, which is interpretable, effectively forecasted the likelihood of MOF in individuals suffering from acute pancreatitis, showcasing strong predictive performance in both internal and external validation groups. Additionally, this model has been implemented as an online tool for risk assessment to improve its practical application in clinical settings.

Authors

Institutions

Publication Details

Journal
European journal of medical research
Published
2026-10-06
DOI
https://doi.org/10.1186/s40001-026-05298-5
Primary Topic
Pancreatitis Pathology and Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Development and validation of an explainable machine learning model for predicting multiple organ failure in patients with acute pancreatitis: a multicenter cohort study

Qinyang Du, Rongshen Guan, Yi Hao, Gaopeng Li et al.
European journal of medical research
Pancreatitis Pathology and Treatment
article

Development and validation of an explainable machine learning model for predicting multiple organ failure in patients with acute pancreatitis: a multicenter cohort study

Qinyang Du, Rongshen Guan, Yi Hao, Gaopeng Li, Yi Wang, Qirui Zhang
article en

Abstract

Acute pancreatitis can lead to a serious and life-threatening situation known as multiple organ failure (MOF). Timely and precise forecasting and detection of MOF are essential. Predictive models utilizing machine learning have shown potential in forecasting MOF in emergency situations. Nevertheless, their use for predicting MOF specifically in patients with acute pancreatitis is still not widespread. This research seeks to create and confirm a machine learning-based model for predicting MOF in individuals suffering from acute pancreatitis. This research utilized two retrospective cohorts for the purposes of developing, validating, and testing the model. The derivation and validation cohorts were sourced from the MIMIC-IV database, which was divided randomly into two segments (70% for model development and 30% for internal validation). For external validation, a retrospective cohort from The Department of Hepatobiliary and Pancreatic Surgery at the Third Hospital of Shanxi Medical University (DHPS cohort) was used. After performing several data preprocessing techniques, including interpolation and standardization, we applied four methods for feature selection, ultimately identifying 15 key features. Seven machine learning algorithms were utilized to create predictive models, with their performance assessed through various metrics such as ROC (receiver operating characteristic), DCA (decision curve analysis), PRC (precision recall curve), calibration curve, and confusion matrice. The final model's interpretation was carried out using SHAP technology, and a web-based risk calculator was created for use in clinical settings. Finally, the final model was used to compare with the SOFA score and APACHE Ⅱ score, respectively, to determine its clinical value. The model was created utilizing data from the MIMIC-Ⅳ database (n=582) and underwent validation and testing with both the MIMIC-Ⅳ (n=250) and The Department of Hepatobiliary and Pancreatic Surgery at the Third Hospital of Shanxi Medical University (n=474) datasets. We identified the best feature combination through four different selection techniques and seven machine learning algorithms. A variable that was recognized by all four selection methods was included in the model development process. The most effective model, CatBoost, was built using 15 easily accessible admission features, including INR (international normalized ratio), Cr (creatinine), PO2 (partial pressure of oxygen), Nbps (non-invasive blood pressure systolic), BIL (bilirubin), HTN (hypertension), PH (potential of hydrogen), LAC (lactate), Hb (hemoglobin), CKD (chronic kidney disease), T (temperature), ALT (alanine aminotransferase), HR (heart rate), Na + (sodion), ABX (antibiotic). During validation, it achieved an AUC value of 0.939, an F1 score of 0.795, and an accuracy rate of 0.856. Subsequently, hyperparameter optimization was conducted on the derivation cohort via 5-fold cross-validation and grid search. The final model was assessed on the test dataset, yielding an AUC of 0.855, an F1 score of 0.658, and an accuracy of 0.761. Additionally, SHAP analysis indicated that INR, Cr, and PO2 are the three most significant variables affecting the model's predictions. This model has been transformed into an online clinical tool to enhance its application in healthcare environments. In validation, CatBoost demonstrated superior discrimination (AUC: 0.941, 95%CI: 0.914–0.968) compared to both APACHE II (0.802, 95%CI: 0.747–0.858) and SOFA (0.819, 95%CI: 0.765–0.874) through ROC and decision curve analyses. The CatBoost model, which is interpretable, effectively forecasted the likelihood of MOF in individuals suffering from acute pancreatitis, showcasing strong predictive performance in both internal and external validation groups. Additionally, this model has been implemented as an online tool for risk assessment to improve its practical application in clinical settings.

European journal of medical research
Shanxi Medical University (CN), Shanxi Academy of Medical Sciences (CN)
Openalex Percentile: Top 9%
Pancreatitis Pathology and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.