Inter-Model Variability of ChatGPT, Gemini and Claude in Treatment Recommendations for Unruptured Intracranial Aneurysms
Abstract Patients increasingly consult large language models (LLMs) before specialist review, yet whether frontier models agree with one another in nuanced clinical domains such as unruptured intracranial aneurysm (UIA) management remains uncharacterised. We therefore quantified inter-model variability across ChatGPT, Gemini and Claude, anchored against neurovascular multidisciplinary team (MDT) consensus and the Unruptured Intracranial Aneurysm Treatment Score (UIATS). Sixty-seven UIA cases referred to our neurovascular service (January–December 2025) were retrospectively analysed. De-identified clinical vignettes were submitted to Claude Opus 4.6, ChatGPT-5.4 and Gemini 3 Pro Thinking, each run five times. Within-model reproducibility was assessed using Fleiss’ κ; inter-model agreement via Cohen’s κ and McNemar’s test; and each LLM’s majority-vote anchored against MDT and UIATS using the same methods. Within-model reproducibility was almost perfect (Fleiss’ κ 0.837–0.860). Pairwise inter-model agreement was asymmetric: ChatGPT–Gemini behaved near-identically (Cohen’s κ = 0.850, 95% CI 0.71–0.97), whereas Claude diverged from both (κ = 0.688 and 0.667). Recommendations were non-unanimous in 13/67 cases (19.4%); Claude was the sole outlier in 8/13 (conservative in 7). Gemini was significantly more pro-treatment than Claude (McNemar p = 0.0117). Against MDT, Gemini and ChatGPT showed significant pro-treatment propensity ( p = 0.0022, p = 0.0153). Claude was significantly more conservative than UIATS ( p = 0.0162). Frontier LLMs are highly reproducible internally but diverge from one another asymmetrically. ChatGPT and Gemini behave near-identically while Claude diverges conservatively. Clinicians should anticipate AI-driven treatment expectations and future work should explore prompting, patient sentiment, and multicentre moderators.
Authors
- Calvin Gan (ORCID: https://orcid.org/0000-0001-8766-240X)
- Lee‐Anne Slater (ORCID: https://orcid.org/0000-0002-1140-8664)
- Ronil Chandra
- Anish Narayan
- Andrew Gauden
- Malik Farooq
- Frederick Mariajoseph
- Idrees Sher
- Adrian Praeger
- Justin Moore
- Hamed Asadi
Institutions
- Monash Health (AU)
- Austin Health (AU)
- Monash University (AU)
Publication Details
- Journal
- Neurosurgical Review
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1007/s10143-026-04498-1
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00