Benchmarking advanced large language models for refractive surgery recommendation: a multi-model, real-world evaluation
To comprehensively evaluate the performance of multiple advanced large language models (LLMs) in simulating real-world clinical decision-making for refractive surgery, comparing their accuracy, consistency, and cost-effectiveness against an intermediate-level ophthalmologist. This retrospective study analyzed preoperative data from 11,966 consecutive patients. A gold standard was established by a panel of three senior refractive surgeons (Inter-expert agreement was excellent; kappa > 0.85) who provided recommendation scores (0-100) and suitability classifications for four procedures: Femtosecond LASIK, SMILE, TransPRK, and ICL. Five LLMs (DeepSeek-Chat, GLM-4.7, GPT-4o, Kimi-K2-Thinking, Qwen-Max) and one intermediate physician independently evaluated all cases via a structured expert-mimicking prompt. Performance was assessed using accuracy, AUC, Cohen’s kappa, correlation coefficients (R), and regression errors (RMSE, MAE). Response time and cost per query were also analyzed. LLMs demonstrated superior or comparable performance to the intermediate physician across tasks. In binary classification, top-performing LLMs achieved accuracies >98.5% and AUCs >0.96 for LASIK and SMILE. For multi-class agreement, Qwen-Max showed the highest consistency with experts (kappa up to 0.743). However, multi-class consistency (Cohen’s kappa) was more modest, indicating LLMs remain best suited as decision-support tools. In regression tasks, Qwen-Max and DeepSeek-Chat exhibited strong correlation with expert scores. The intermediate physician showed significantly lower performance, particularly in complex classifications. DeepSeek-Chat offered the best cost-efficiency with lowest cost and fastest speed, while GPT-4o was the most expensive. Advanced LLMs show strong potential as clinical decision-support tools in refractive surgery planning, with top-performing models approaching intermediate-level physician performance in binary classification tasks. However, their multi-class agreement remains moderate, and they should be positioned as assistive tools rather than autonomous decision-makers. Cost and efficiency advantages make them particularly suitable for large-scale screening and resource-limited settings.
Authors
- Ran Wei (ORCID: https://orcid.org/0000-0003-4293-6686)
- Qi Wan (ORCID: https://orcid.org/0000-0002-6978-4777)
- Ying-ping Deng
- Ke Ma
- Jing Tang
Institutions
- Sichuan University (CN)
- West China Hospital of Sichuan University (CN)
Publication Details
- Journal
- BMC Medicine
- Published
- 2026-09-25
- DOI
- https://doi.org/10.1186/s12916-026-05262-4
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00