Diagnostic performance of ChatGPT for acute appendicitis compared with the Alvarado score: a retrospective diagnostic accuracy study
Acute appendicitis remains the most common cause of acute abdomen and is associated with serious complications such as perforation if not diagnosed promptly. While laboratory tests and clinical findings are essential, their interpretation can be challenging and sometimes misleading. Artificial intelligence (AI), particularly large language models (LLMs) like ChatGPT, has shown promise in supporting clinical decision-making and enhancing diagnostic accuracy. This study aimed to evaluate the ability of ChatGPT-5 to predict the likelihood of appendicitis using clinical data, with diagnoses confirmed by histopathology and comparing it to Alvarado score which is most commonly known clinical scoring system to diagnose acute appendicitis. This retrospective comparative study evaluated the performance of ChatGPT-5 in predicting acute appendicitis across two phases. In Phase 1, the model was tested without prior context learning (i.e., without exposure to any previously labeled cases) to simulate real-world use of readily available LLMs. In Phase 2, it was re-tested after exposure to confirmed case outcomes. A structured prompting strategy was applied, incorporating patient demographics, clinical symptoms, laboratory data, and clinical context, with fixed output constraints across all cases. The dataset was divided into 70% (114 cases) for Phase 1 evaluation and 30% (48 cases) for Phase 2 validation. To address potential order leakage- a bias arising when predictions depend on case sequence rather than clinical reasoning- the cases were assessed both with and without randomization. Paired comparisons of sensitivity, specificity, and accuracy between ChatGPT-5 and the Alvarado score were performed using McNemar’s test, and Cohen’s kappa was used to assess agreement. In Phase 1, ChatGPT-5 without randomization achieved a sensitivity of 100%, specificity of 81.6%, positive predictive value (PPV) of 91.6%, negative predictive value (NPV) of 100%, and overall accuracy of 93.9%, compared with the Alvarado score (sensitivity 67.1%, specificity 92.1%, PPV 94.4%, NPV 58.3%, accuracy 75.4%). In Phase 2, ChatGPT-5 without randomization reached a sensitivity of 100%, specificity of 93.8%, PPV of 97.0%, NPV of 100%, and accuracy of 97.9%, comparable to the Alvarado score (sensitivity 96.9%, specificity 93.8%, PPV 96.9%, NPV 93.8%, accuracy 95.8%). Within this retrospective, surgically confirmed cohort with a high clinical suspicion of acute appendicitis, ChatGPT-5 showed high diagnostic sensitivity for excluding acute appendicitis when guided by an optimized prompt, even without prior context learning. Proper randomization of prompt data appears essential to minimize order leakage and to provide a more conservative and reliable estimate of model performance. ChatGPT-5 cannot yet be relied upon to confirm acute appendicitis, given the higher specificity of the Alvarado score, and its diagnostic sensitivity remained largely unaffected by context learning, with only a modest, non-significant gain in specificity. These findings should be regarded as hypothesis-generating and require prospective validation in a broader emergency-department population before any clinical implementation.
Authors
- Amr A. Elgharib (ORCID: https://orcid.org/0000-0001-7938-8316)
- Ayman Shemes (ORCID: https://orcid.org/0000-0002-6824-7865)
- Nour Elwakeel
- Rawan S. Ibrahim
- Moustafa I. Ghanem
Institutions
- Mansoura University (EG)
- Mansoura University Hospital (EG)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1007/s44163-026-02332-7
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00