Diagnostic performance of ChatGPT for acute appendicitis compared with the Alvarado score: a retrospective diagnostic accuracy study

Acute appendicitis remains the most common cause of acute abdomen and is associated with serious complications such as perforation if not diagnosed promptly. While laboratory tests and clinical findings are essential, their interpretation can be challenging and sometimes misleading. Artificial intelligence (AI), particularly large language models (LLMs) like ChatGPT, has shown promise in supporting clinical decision-making and enhancing diagnostic accuracy. This study aimed to evaluate the ability of ChatGPT-5 to predict the likelihood of appendicitis using clinical data, with diagnoses confirmed by histopathology and comparing it to Alvarado score which is most commonly known clinical scoring system to diagnose acute appendicitis. This retrospective comparative study evaluated the performance of ChatGPT-5 in predicting acute appendicitis across two phases. In Phase 1, the model was tested without prior context learning (i.e., without exposure to any previously labeled cases) to simulate real-world use of readily available LLMs. In Phase 2, it was re-tested after exposure to confirmed case outcomes. A structured prompting strategy was applied, incorporating patient demographics, clinical symptoms, laboratory data, and clinical context, with fixed output constraints across all cases. The dataset was divided into 70% (114 cases) for Phase 1 evaluation and 30% (48 cases) for Phase 2 validation. To address potential order leakage- a bias arising when predictions depend on case sequence rather than clinical reasoning- the cases were assessed both with and without randomization. Paired comparisons of sensitivity, specificity, and accuracy between ChatGPT-5 and the Alvarado score were performed using McNemar’s test, and Cohen’s kappa was used to assess agreement. In Phase 1, ChatGPT-5 without randomization achieved a sensitivity of 100%, specificity of 81.6%, positive predictive value (PPV) of 91.6%, negative predictive value (NPV) of 100%, and overall accuracy of 93.9%, compared with the Alvarado score (sensitivity 67.1%, specificity 92.1%, PPV 94.4%, NPV 58.3%, accuracy 75.4%). In Phase 2, ChatGPT-5 without randomization reached a sensitivity of 100%, specificity of 93.8%, PPV of 97.0%, NPV of 100%, and accuracy of 97.9%, comparable to the Alvarado score (sensitivity 96.9%, specificity 93.8%, PPV 96.9%, NPV 93.8%, accuracy 95.8%). Within this retrospective, surgically confirmed cohort with a high clinical suspicion of acute appendicitis, ChatGPT-5 showed high diagnostic sensitivity for excluding acute appendicitis when guided by an optimized prompt, even without prior context learning. Proper randomization of prompt data appears essential to minimize order leakage and to provide a more conservative and reliable estimate of model performance. ChatGPT-5 cannot yet be relied upon to confirm acute appendicitis, given the higher specificity of the Alvarado score, and its diagnostic sensitivity remained largely unaffected by context learning, with only a modest, non-significant gain in specificity. These findings should be regarded as hypothesis-generating and require prospective validation in a broader emergency-department population before any clinical implementation.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-10-05
DOI
https://doi.org/10.1007/s44163-026-02332-7
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Diagnostic performance of ChatGPT for acute appendicitis compared with the Alvarado score: a retrospective diagnostic accuracy study

Amr A. Elgharib, Ayman Shemes, Nour Elwakeel, Rawan S. Ibrahim et al.
Discover Artificial Intelligence
Artificial Intelligence in Healthcare and Education
article

Diagnostic performance of ChatGPT for acute appendicitis compared with the Alvarado score: a retrospective diagnostic accuracy study

Amr A. Elgharib, Ayman Shemes, Nour Elwakeel, Rawan S. Ibrahim, Moustafa I. Ghanem
article en

Abstract

Acute appendicitis remains the most common cause of acute abdomen and is associated with serious complications such as perforation if not diagnosed promptly. While laboratory tests and clinical findings are essential, their interpretation can be challenging and sometimes misleading. Artificial intelligence (AI), particularly large language models (LLMs) like ChatGPT, has shown promise in supporting clinical decision-making and enhancing diagnostic accuracy. This study aimed to evaluate the ability of ChatGPT-5 to predict the likelihood of appendicitis using clinical data, with diagnoses confirmed by histopathology and comparing it to Alvarado score which is most commonly known clinical scoring system to diagnose acute appendicitis. This retrospective comparative study evaluated the performance of ChatGPT-5 in predicting acute appendicitis across two phases. In Phase 1, the model was tested without prior context learning (i.e., without exposure to any previously labeled cases) to simulate real-world use of readily available LLMs. In Phase 2, it was re-tested after exposure to confirmed case outcomes. A structured prompting strategy was applied, incorporating patient demographics, clinical symptoms, laboratory data, and clinical context, with fixed output constraints across all cases. The dataset was divided into 70% (114 cases) for Phase 1 evaluation and 30% (48 cases) for Phase 2 validation. To address potential order leakage- a bias arising when predictions depend on case sequence rather than clinical reasoning- the cases were assessed both with and without randomization. Paired comparisons of sensitivity, specificity, and accuracy between ChatGPT-5 and the Alvarado score were performed using McNemar’s test, and Cohen’s kappa was used to assess agreement. In Phase 1, ChatGPT-5 without randomization achieved a sensitivity of 100%, specificity of 81.6%, positive predictive value (PPV) of 91.6%, negative predictive value (NPV) of 100%, and overall accuracy of 93.9%, compared with the Alvarado score (sensitivity 67.1%, specificity 92.1%, PPV 94.4%, NPV 58.3%, accuracy 75.4%). In Phase 2, ChatGPT-5 without randomization reached a sensitivity of 100%, specificity of 93.8%, PPV of 97.0%, NPV of 100%, and accuracy of 97.9%, comparable to the Alvarado score (sensitivity 96.9%, specificity 93.8%, PPV 96.9%, NPV 93.8%, accuracy 95.8%). Within this retrospective, surgically confirmed cohort with a high clinical suspicion of acute appendicitis, ChatGPT-5 showed high diagnostic sensitivity for excluding acute appendicitis when guided by an optimized prompt, even without prior context learning. Proper randomization of prompt data appears essential to minimize order leakage and to provide a more conservative and reliable estimate of model performance. ChatGPT-5 cannot yet be relied upon to confirm acute appendicitis, given the higher specificity of the Alvarado score, and its diagnostic sensitivity remained largely unaffected by context learning, with only a modest, non-significant gain in specificity. These findings should be regarded as hypothesis-generating and require prospective validation in a broader emergency-department population before any clinical implementation.

Discover Artificial IntelligenceVol. 6(1)
Mansoura University (EG), Mansoura University Hospital (EG)
Openalex Percentile: Top 18%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.