Inter-model agreement and determinants of recommendation discordance in large language model–assisted refractive surgery planning: a comparative study of ChatGPT and Grok
To evaluate the consistency of refractive surgery recommendations generated by two multimodal large language models (LLMs), ChatGPT and Grok, using identical corneal tomography data, and to identify clinical and tomographic factors associated with inter-model disagreement. This observational study included 69 right eyes of patients who underwent photorefractive keratectomy (PRK) or transepithelial PRK for myopia and/or astigmatism. Pentacam four-map displays, Belin/Ambrosio Enhanced Ectasia Display images, patient demographics, and refractive errors were independently submitted to ChatGPT and Grok using an identical standardized prompt requesting ranked refractive surgery recommendations. Agreement was assessed using Cohen’s kappa (κ). Logistic regression analyses were used to identify factors associated with disagreement, while receiver operating characteristic (ROC) analysis and pairwise DeLong tests were used to evaluate the discriminatory performance of significant tomographic parameters. Both LLMs most frequently recommended SMILE as the first-line procedure (ChatGPT, 63.8%; Grok 60.9%). Agreement between the models was slight to fair for first- (κ = 0.128), second- (κ = 0.233), and third-ranked (κ = 0.238) recommendations, with complete agreement observed in 24 of 69 eyes (34.8%). Lower pachymetry and higher Dp, Dt, and BAD-D values were significantly associated with disagreement in the univariate analysis; however, none remained independently associated after multivariate adjustment. ROC analysis demonstrated significant discriminatory performance for pachymetry, Dp, Dt, and BAD-D, with BAD-D having the highest AUC (0.717). Pairwise DeLong comparisons revealed no significant differences between the ROC curves. ChatGPT and Grok demonstrated limited agreement in refractive surgery recommendations despite identical clinical and tomographic inputs. Corneal thickness and BAD-derived tomographic indices were associated with greater inter-model disagreement, although no independent predictors were identified. These findings demonstrate the current variability among LLMs and support their use as decision-support tools that complement, rather than replace, expert clinical judgment. Disagreements between independent LLMs may provide clinically useful information.
Authors
- Konuralp Yakar (ORCID: https://orcid.org/0000-0002-3839-5699)
Institutions
- Samsun University (TR)
- Liv Hospital (TR)
Publication Details
- Journal
- BMC Ophthalmology
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1186/s12886-026-05403-6
- Primary Topic
- Corneal surgery and disorders
- Type
- article
- Field-Weighted Citation Impact
- 0.00