Inter-model agreement and determinants of recommendation discordance in large language model–assisted refractive surgery planning: a comparative study of ChatGPT and Grok

To evaluate the consistency of refractive surgery recommendations generated by two multimodal large language models (LLMs), ChatGPT and Grok, using identical corneal tomography data, and to identify clinical and tomographic factors associated with inter-model disagreement. This observational study included 69 right eyes of patients who underwent photorefractive keratectomy (PRK) or transepithelial PRK for myopia and/or astigmatism. Pentacam four-map displays, Belin/Ambrosio Enhanced Ectasia Display images, patient demographics, and refractive errors were independently submitted to ChatGPT and Grok using an identical standardized prompt requesting ranked refractive surgery recommendations. Agreement was assessed using Cohen’s kappa (κ). Logistic regression analyses were used to identify factors associated with disagreement, while receiver operating characteristic (ROC) analysis and pairwise DeLong tests were used to evaluate the discriminatory performance of significant tomographic parameters. Both LLMs most frequently recommended SMILE as the first-line procedure (ChatGPT, 63.8%; Grok 60.9%). Agreement between the models was slight to fair for first- (κ = 0.128), second- (κ = 0.233), and third-ranked (κ = 0.238) recommendations, with complete agreement observed in 24 of 69 eyes (34.8%). Lower pachymetry and higher Dp, Dt, and BAD-D values were significantly associated with disagreement in the univariate analysis; however, none remained independently associated after multivariate adjustment. ROC analysis demonstrated significant discriminatory performance for pachymetry, Dp, Dt, and BAD-D, with BAD-D having the highest AUC (0.717). Pairwise DeLong comparisons revealed no significant differences between the ROC curves. ChatGPT and Grok demonstrated limited agreement in refractive surgery recommendations despite identical clinical and tomographic inputs. Corneal thickness and BAD-derived tomographic indices were associated with greater inter-model disagreement, although no independent predictors were identified. These findings demonstrate the current variability among LLMs and support their use as decision-support tools that complement, rather than replace, expert clinical judgment. Disagreements between independent LLMs may provide clinically useful information.

Authors

Institutions

Publication Details

Journal
BMC Ophthalmology
Published
2026-09-30
DOI
https://doi.org/10.1186/s12886-026-05403-6
Primary Topic
Corneal surgery and disorders
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Inter-model agreement and determinants of recommendation discordance in large language model–assisted refractive surgery planning: a comparative study of ChatGPT and Grok

Konuralp Yakar
BMC Ophthalmology
Corneal surgery and disorders
article

Inter-model agreement and determinants of recommendation discordance in large language model–assisted refractive surgery planning: a comparative study of ChatGPT and Grok

Konuralp Yakar
article en

Abstract

To evaluate the consistency of refractive surgery recommendations generated by two multimodal large language models (LLMs), ChatGPT and Grok, using identical corneal tomography data, and to identify clinical and tomographic factors associated with inter-model disagreement. This observational study included 69 right eyes of patients who underwent photorefractive keratectomy (PRK) or transepithelial PRK for myopia and/or astigmatism. Pentacam four-map displays, Belin/Ambrosio Enhanced Ectasia Display images, patient demographics, and refractive errors were independently submitted to ChatGPT and Grok using an identical standardized prompt requesting ranked refractive surgery recommendations. Agreement was assessed using Cohen’s kappa (κ). Logistic regression analyses were used to identify factors associated with disagreement, while receiver operating characteristic (ROC) analysis and pairwise DeLong tests were used to evaluate the discriminatory performance of significant tomographic parameters. Both LLMs most frequently recommended SMILE as the first-line procedure (ChatGPT, 63.8%; Grok 60.9%). Agreement between the models was slight to fair for first- (κ = 0.128), second- (κ = 0.233), and third-ranked (κ = 0.238) recommendations, with complete agreement observed in 24 of 69 eyes (34.8%). Lower pachymetry and higher Dp, Dt, and BAD-D values were significantly associated with disagreement in the univariate analysis; however, none remained independently associated after multivariate adjustment. ROC analysis demonstrated significant discriminatory performance for pachymetry, Dp, Dt, and BAD-D, with BAD-D having the highest AUC (0.717). Pairwise DeLong comparisons revealed no significant differences between the ROC curves. ChatGPT and Grok demonstrated limited agreement in refractive surgery recommendations despite identical clinical and tomographic inputs. Corneal thickness and BAD-derived tomographic indices were associated with greater inter-model disagreement, although no independent predictors were identified. These findings demonstrate the current variability among LLMs and support their use as decision-support tools that complement, rather than replace, expert clinical judgment. Disagreements between independent LLMs may provide clinically useful information.

BMC Ophthalmology
Samsun University (TR), Liv Hospital (TR)
Reduced inequalities
Openalex Percentile: Top 12%
Corneal surgery and disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.