Inter-Model Variability of ChatGPT, Gemini and Claude in Treatment Recommendations for Unruptured Intracranial Aneurysms

Abstract Patients increasingly consult large language models (LLMs) before specialist review, yet whether frontier models agree with one another in nuanced clinical domains such as unruptured intracranial aneurysm (UIA) management remains uncharacterised. We therefore quantified inter-model variability across ChatGPT, Gemini and Claude, anchored against neurovascular multidisciplinary team (MDT) consensus and the Unruptured Intracranial Aneurysm Treatment Score (UIATS). Sixty-seven UIA cases referred to our neurovascular service (January–December 2025) were retrospectively analysed. De-identified clinical vignettes were submitted to Claude Opus 4.6, ChatGPT-5.4 and Gemini 3 Pro Thinking, each run five times. Within-model reproducibility was assessed using Fleiss’ κ; inter-model agreement via Cohen’s κ and McNemar’s test; and each LLM’s majority-vote anchored against MDT and UIATS using the same methods. Within-model reproducibility was almost perfect (Fleiss’ κ 0.837–0.860). Pairwise inter-model agreement was asymmetric: ChatGPT–Gemini behaved near-identically (Cohen’s κ = 0.850, 95% CI 0.71–0.97), whereas Claude diverged from both (κ = 0.688 and 0.667). Recommendations were non-unanimous in 13/67 cases (19.4%); Claude was the sole outlier in 8/13 (conservative in 7). Gemini was significantly more pro-treatment than Claude (McNemar p = 0.0117). Against MDT, Gemini and ChatGPT showed significant pro-treatment propensity ( p = 0.0022, p = 0.0153). Claude was significantly more conservative than UIATS ( p = 0.0162). Frontier LLMs are highly reproducible internally but diverge from one another asymmetrically. ChatGPT and Gemini behave near-identically while Claude diverges conservatively. Clinicians should anticipate AI-driven treatment expectations and future work should explore prompting, patient sentiment, and multicentre moderators.

Authors

Institutions

Publication Details

Journal
Neurosurgical Review
Published
2026-09-16
DOI
https://doi.org/10.1007/s10143-026-04498-1
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Inter-Model Variability of ChatGPT, Gemini and Claude in Treatment Recommendations for Unruptured Intracranial Aneurysms

Calvin Gan, Lee‐Anne Slater, Ronil Chandra, Anish Narayan et al.
Neurosurgical Review
Artificial Intelligence in Healthcare and Education
article

Inter-Model Variability of ChatGPT, Gemini and Claude in Treatment Recommendations for Unruptured Intracranial Aneurysms

Calvin Gan, Lee‐Anne Slater, Ronil Chandra, Anish Narayan, Andrew Gauden, Malik Farooq, Frederick Mariajoseph, Idrees Sher, Adrian Praeger, Justin Moore, Hamed Asadi
article en

Abstract

Abstract Patients increasingly consult large language models (LLMs) before specialist review, yet whether frontier models agree with one another in nuanced clinical domains such as unruptured intracranial aneurysm (UIA) management remains uncharacterised. We therefore quantified inter-model variability across ChatGPT, Gemini and Claude, anchored against neurovascular multidisciplinary team (MDT) consensus and the Unruptured Intracranial Aneurysm Treatment Score (UIATS). Sixty-seven UIA cases referred to our neurovascular service (January–December 2025) were retrospectively analysed. De-identified clinical vignettes were submitted to Claude Opus 4.6, ChatGPT-5.4 and Gemini 3 Pro Thinking, each run five times. Within-model reproducibility was assessed using Fleiss’ κ; inter-model agreement via Cohen’s κ and McNemar’s test; and each LLM’s majority-vote anchored against MDT and UIATS using the same methods. Within-model reproducibility was almost perfect (Fleiss’ κ 0.837–0.860). Pairwise inter-model agreement was asymmetric: ChatGPT–Gemini behaved near-identically (Cohen’s κ = 0.850, 95% CI 0.71–0.97), whereas Claude diverged from both (κ = 0.688 and 0.667). Recommendations were non-unanimous in 13/67 cases (19.4%); Claude was the sole outlier in 8/13 (conservative in 7). Gemini was significantly more pro-treatment than Claude (McNemar p = 0.0117). Against MDT, Gemini and ChatGPT showed significant pro-treatment propensity ( p = 0.0022, p = 0.0153). Claude was significantly more conservative than UIATS ( p = 0.0162). Frontier LLMs are highly reproducible internally but diverge from one another asymmetrically. ChatGPT and Gemini behave near-identically while Claude diverges conservatively. Clinicians should anticipate AI-driven treatment expectations and future work should explore prompting, patient sentiment, and multicentre moderators.

Neurosurgical ReviewVol. 49(1)
Monash Health (AU), Austin Health (AU), Monash University (AU)
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.