Machine learning models for predicting active bleeding in patients presenting to the emergency department with upper gastrointestinal bleeding: a retrospective cross-sectional study

Upper gastrointestinal (GI) bleeding is a common emergency department (ED) presentation with substantial morbidity and mortality. Endoscopy is the diagnostic and therapeutic cornerstone, yet pre-endoscopic risk-stratification tools such as the Glasgow–Blatchford Score (GBS) are highly sensitive but suffer from low specificity, leading to a high rate of negative endoscopies. We aimed to develop and internally validate supervised machine-learning (ML) models to predict endoscopically confirmed active bleeding (Forrest Ia/Ib) using routinely available admission data, with the goal of improving specificity and reducing unnecessary endoscopic procedures. A retrospective, observational, cross-sectional, analytical study was conducted in the ED of a tertiary-care academic hospital after ethics committee approval. Patients aged ≥ 18 years who presented between 1 January 2015 and 1 August 2022 with a working diagnosis of upper GI bleeding and underwent endoscopy were screened. Demographics, vital signs, digital rectal examination findings, comorbidities, complete blood count, biochemistry, coagulation, blood gas parameters, GBS, and endoscopic Forrest classification were extracted. Active bleeding was defined as Forrest 1a or 1b. A priori power analysis indicated a minimum target of 1,146 patients (803 training / 343 test); 1,120 candidate cases were identified, of whom 1,009 met inclusion criteria after exclusions — falling short of the training-set target (666 vs. 803) but with the pre-specified validation set of 343 achieved. Of the 1,009 eligible patients, 666 (66.0%) were randomly allocated to the training set and 343 (34.0%) to the internal validation/test set using stratified random sampling in Python v3.9.0. Five ML algorithms — Gaussian Naive Bayes (GNB), K-Nearest Neighbors (KNN), Support Vector Machines (SVM), Random Forest, and Logistic Regression — were trained on admission features and evaluated on the held-out set. Performance was summarised by sensitivity, specificity, positive and negative likelihood ratios, positive and negative predictive values (PPV/NPV), and accuracy. Group comparisons in the test set were performed with Chi-square/Fisher’s exact, independent-samples t and Mann–Whitney U tests in SPSS v24.0 (two-sided p < 0.05). Of 343 patients in the test set, 30 (8.7%) had endoscopically confirmed active bleeding. Compared with the no-active-bleeding group, patients with active bleeding had significantly higher GBS (13.07 ± 1.93 vs. 12.19 ± 2.36; p = 0.039) and lower platelet counts (224.37 ± 95.58 vs. 267.65 ± 118.78 × 10³/µL; p = 0.049). Other admission parameters (age, sex, vital signs, hemoglobin, BUN, INR, pH, lactate) did not differ significantly. Forrest classification in the 153 patients with positive endoscopic findings was Forrest 3 in 58.2%, Forrest 1b in 17.0%, Forrest 1a in 2.6%, Forrest 2a in 8.5%, Forrest 2b in 8.5%, and Forrest 2c in 5.2%. On the held-out test set, Random Forest achieved the highest discrimination (AUC 71.2%; 95% CI 60.4–81.9%; sensitivity 80.0% [95% CI 62.7–90.5%]; NPV 97.0% [95% CI 93.6–98.6%]). Logistic Regression provided a balanced profile (AUC 68.5%; 95% CI 57.6–79.5%; sensitivity 70.0%; NPV 95.9%). The GBS comparator at cutoff ≥ 13 yielded AUC 61.3% (95% CI 50.2–72.4%), sensitivity 66.7%, and NPV 94.4%. RF and LR point estimates exceeded GBS, though confidence intervals overlap and no formal statistical comparison was performed. KNN showed near-zero sensitivity (3.3%) despite high specificity (99.4%), driven by class imbalance. Performance estimates should be interpreted with caution given only 30 outcome events in the test set. In ED patients undergoing endoscopy for suspected upper GI bleeding, ML models trained on routinely collected admission data — particularly Random Forest, Logistic Regression and SVM — achieved high negative predictive values (≥ 95%) for endoscopically confirmed active bleeding. Given the 8.7% prevalence of active bleeding in our cohort, these models showed potentially useful discrimination in this single-centre internal validation cohort and may complement existing risk scores such as the GBS; whether they can support endoscopy triage decisions requires prospective external validation and formal assessment of clinical utility. Prospective multicentre validation is required before clinical deployment.

Authors

Institutions

Publication Details

Journal
BMC Emergency Medicine
Published
2026-09-11
DOI
https://doi.org/10.1186/s12873-026-01772-9
Primary Topic
Gastrointestinal Bleeding Diagnosis and Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Machine learning models for predicting active bleeding in patients presenting to the emergency department with upper gastrointestinal bleeding: a retrospective cross-sectional study

Bilal Yeniyurt, Serkan Doğan, Utku Murat Kalafat, Melih Uçan et al.
BMC Emergency Medicine
Gastrointestinal Bleeding Diagnosis and Treatment
article

Machine learning models for predicting active bleeding in patients presenting to the emergency department with upper gastrointestinal bleeding: a retrospective cross-sectional study

Bilal Yeniyurt, Serkan Doğan, Utku Murat Kalafat, Melih Uçan, Ramiz Yazıcı, Salih Fettahoğlu, Burçe Serra Koçkan, Süreyya Tuba Fettahoğlu, Ahmed Edizer
article en

Abstract

Upper gastrointestinal (GI) bleeding is a common emergency department (ED) presentation with substantial morbidity and mortality. Endoscopy is the diagnostic and therapeutic cornerstone, yet pre-endoscopic risk-stratification tools such as the Glasgow–Blatchford Score (GBS) are highly sensitive but suffer from low specificity, leading to a high rate of negative endoscopies. We aimed to develop and internally validate supervised machine-learning (ML) models to predict endoscopically confirmed active bleeding (Forrest Ia/Ib) using routinely available admission data, with the goal of improving specificity and reducing unnecessary endoscopic procedures. A retrospective, observational, cross-sectional, analytical study was conducted in the ED of a tertiary-care academic hospital after ethics committee approval. Patients aged ≥ 18 years who presented between 1 January 2015 and 1 August 2022 with a working diagnosis of upper GI bleeding and underwent endoscopy were screened. Demographics, vital signs, digital rectal examination findings, comorbidities, complete blood count, biochemistry, coagulation, blood gas parameters, GBS, and endoscopic Forrest classification were extracted. Active bleeding was defined as Forrest 1a or 1b. A priori power analysis indicated a minimum target of 1,146 patients (803 training / 343 test); 1,120 candidate cases were identified, of whom 1,009 met inclusion criteria after exclusions — falling short of the training-set target (666 vs. 803) but with the pre-specified validation set of 343 achieved. Of the 1,009 eligible patients, 666 (66.0%) were randomly allocated to the training set and 343 (34.0%) to the internal validation/test set using stratified random sampling in Python v3.9.0. Five ML algorithms — Gaussian Naive Bayes (GNB), K-Nearest Neighbors (KNN), Support Vector Machines (SVM), Random Forest, and Logistic Regression — were trained on admission features and evaluated on the held-out set. Performance was summarised by sensitivity, specificity, positive and negative likelihood ratios, positive and negative predictive values (PPV/NPV), and accuracy. Group comparisons in the test set were performed with Chi-square/Fisher’s exact, independent-samples t and Mann–Whitney U tests in SPSS v24.0 (two-sided p < 0.05). Of 343 patients in the test set, 30 (8.7%) had endoscopically confirmed active bleeding. Compared with the no-active-bleeding group, patients with active bleeding had significantly higher GBS (13.07 ± 1.93 vs. 12.19 ± 2.36; p = 0.039) and lower platelet counts (224.37 ± 95.58 vs. 267.65 ± 118.78 × 10³/µL; p = 0.049). Other admission parameters (age, sex, vital signs, hemoglobin, BUN, INR, pH, lactate) did not differ significantly. Forrest classification in the 153 patients with positive endoscopic findings was Forrest 3 in 58.2%, Forrest 1b in 17.0%, Forrest 1a in 2.6%, Forrest 2a in 8.5%, Forrest 2b in 8.5%, and Forrest 2c in 5.2%. On the held-out test set, Random Forest achieved the highest discrimination (AUC 71.2%; 95% CI 60.4–81.9%; sensitivity 80.0% [95% CI 62.7–90.5%]; NPV 97.0% [95% CI 93.6–98.6%]). Logistic Regression provided a balanced profile (AUC 68.5%; 95% CI 57.6–79.5%; sensitivity 70.0%; NPV 95.9%). The GBS comparator at cutoff ≥ 13 yielded AUC 61.3% (95% CI 50.2–72.4%), sensitivity 66.7%, and NPV 94.4%. RF and LR point estimates exceeded GBS, though confidence intervals overlap and no formal statistical comparison was performed. KNN showed near-zero sensitivity (3.3%) despite high specificity (99.4%), driven by class imbalance. Performance estimates should be interpreted with caution given only 30 outcome events in the test set. In ED patients undergoing endoscopy for suspected upper GI bleeding, ML models trained on routinely collected admission data — particularly Random Forest, Logistic Regression and SVM — achieved high negative predictive values (≥ 95%) for endoscopically confirmed active bleeding. Given the 8.7% prevalence of active bleeding in our cohort, these models showed potentially useful discrimination in this single-centre internal validation cohort and may complement existing risk scores such as the GBS; whether they can support endoscopy triage decisions requires prospective external validation and formal assessment of clinical utility. Prospective multicentre validation is required before clinical deployment.

BMC Emergency Medicine
Sivas State Hospital (TR), Sağlık Bilimleri Üniversitesi (TR), İstanbul Kanuni Sultan Süleyman Eğitim ve Araştırma Hastanesi (TR), Marmara University (TR)
Good health and well-being
Openalex Percentile: Top 9%
Gastrointestinal Bleeding Diagnosis and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.