Machine learning models for predicting active bleeding in patients presenting to the emergency department with upper gastrointestinal bleeding: a retrospective cross-sectional study
Upper gastrointestinal (GI) bleeding is a common emergency department (ED) presentation with substantial morbidity and mortality. Endoscopy is the diagnostic and therapeutic cornerstone, yet pre-endoscopic risk-stratification tools such as the Glasgow–Blatchford Score (GBS) are highly sensitive but suffer from low specificity, leading to a high rate of negative endoscopies. We aimed to develop and internally validate supervised machine-learning (ML) models to predict endoscopically confirmed active bleeding (Forrest Ia/Ib) using routinely available admission data, with the goal of improving specificity and reducing unnecessary endoscopic procedures. A retrospective, observational, cross-sectional, analytical study was conducted in the ED of a tertiary-care academic hospital after ethics committee approval. Patients aged ≥ 18 years who presented between 1 January 2015 and 1 August 2022 with a working diagnosis of upper GI bleeding and underwent endoscopy were screened. Demographics, vital signs, digital rectal examination findings, comorbidities, complete blood count, biochemistry, coagulation, blood gas parameters, GBS, and endoscopic Forrest classification were extracted. Active bleeding was defined as Forrest 1a or 1b. A priori power analysis indicated a minimum target of 1,146 patients (803 training / 343 test); 1,120 candidate cases were identified, of whom 1,009 met inclusion criteria after exclusions — falling short of the training-set target (666 vs. 803) but with the pre-specified validation set of 343 achieved. Of the 1,009 eligible patients, 666 (66.0%) were randomly allocated to the training set and 343 (34.0%) to the internal validation/test set using stratified random sampling in Python v3.9.0. Five ML algorithms — Gaussian Naive Bayes (GNB), K-Nearest Neighbors (KNN), Support Vector Machines (SVM), Random Forest, and Logistic Regression — were trained on admission features and evaluated on the held-out set. Performance was summarised by sensitivity, specificity, positive and negative likelihood ratios, positive and negative predictive values (PPV/NPV), and accuracy. Group comparisons in the test set were performed with Chi-square/Fisher’s exact, independent-samples t and Mann–Whitney U tests in SPSS v24.0 (two-sided p < 0.05). Of 343 patients in the test set, 30 (8.7%) had endoscopically confirmed active bleeding. Compared with the no-active-bleeding group, patients with active bleeding had significantly higher GBS (13.07 ± 1.93 vs. 12.19 ± 2.36; p = 0.039) and lower platelet counts (224.37 ± 95.58 vs. 267.65 ± 118.78 × 10³/µL; p = 0.049). Other admission parameters (age, sex, vital signs, hemoglobin, BUN, INR, pH, lactate) did not differ significantly. Forrest classification in the 153 patients with positive endoscopic findings was Forrest 3 in 58.2%, Forrest 1b in 17.0%, Forrest 1a in 2.6%, Forrest 2a in 8.5%, Forrest 2b in 8.5%, and Forrest 2c in 5.2%. On the held-out test set, Random Forest achieved the highest discrimination (AUC 71.2%; 95% CI 60.4–81.9%; sensitivity 80.0% [95% CI 62.7–90.5%]; NPV 97.0% [95% CI 93.6–98.6%]). Logistic Regression provided a balanced profile (AUC 68.5%; 95% CI 57.6–79.5%; sensitivity 70.0%; NPV 95.9%). The GBS comparator at cutoff ≥ 13 yielded AUC 61.3% (95% CI 50.2–72.4%), sensitivity 66.7%, and NPV 94.4%. RF and LR point estimates exceeded GBS, though confidence intervals overlap and no formal statistical comparison was performed. KNN showed near-zero sensitivity (3.3%) despite high specificity (99.4%), driven by class imbalance. Performance estimates should be interpreted with caution given only 30 outcome events in the test set. In ED patients undergoing endoscopy for suspected upper GI bleeding, ML models trained on routinely collected admission data — particularly Random Forest, Logistic Regression and SVM — achieved high negative predictive values (≥ 95%) for endoscopically confirmed active bleeding. Given the 8.7% prevalence of active bleeding in our cohort, these models showed potentially useful discrimination in this single-centre internal validation cohort and may complement existing risk scores such as the GBS; whether they can support endoscopy triage decisions requires prospective external validation and formal assessment of clinical utility. Prospective multicentre validation is required before clinical deployment.
Authors
- Bilal Yeniyurt (ORCID: https://orcid.org/0000-0001-9881-3437)
- Serkan Doğan (ORCID: https://orcid.org/0000-0001-8923-2489)
- Utku Murat Kalafat (ORCID: https://orcid.org/0000-0003-1749-8098)
- Melih Uçan (ORCID: https://orcid.org/0000-0001-6826-1053)
- Ramiz Yazıcı (ORCID: https://orcid.org/0000-0001-9210-914X)
- Salih Fettahoğlu (ORCID: https://orcid.org/0000-0003-0405-3938)
- Burçe Serra Koçkan (ORCID: https://orcid.org/0000-0001-8996-2959)
- Süreyya Tuba Fettahoğlu (ORCID: https://orcid.org/0000-0001-7882-5086)
- Ahmed Edizer
Institutions
- Sivas State Hospital (TR)
- Sağlık Bilimleri Üniversitesi (TR)
- İstanbul Kanuni Sultan Süleyman Eğitim ve Araştırma Hastanesi (TR)
- Marmara University (TR)
Publication Details
- Journal
- BMC Emergency Medicine
- Published
- 2026-09-11
- DOI
- https://doi.org/10.1186/s12873-026-01772-9
- Primary Topic
- Gastrointestinal Bleeding Diagnosis and Treatment
- Type
- article
- Field-Weighted Citation Impact
- 0.00