Benchmarking advanced large language models for refractive surgery recommendation: a multi-model, real-world evaluation

To comprehensively evaluate the performance of multiple advanced large language models (LLMs) in simulating real-world clinical decision-making for refractive surgery, comparing their accuracy, consistency, and cost-effectiveness against an intermediate-level ophthalmologist. This retrospective study analyzed preoperative data from 11,966 consecutive patients. A gold standard was established by a panel of three senior refractive surgeons (Inter-expert agreement was excellent; kappa > 0.85) who provided recommendation scores (0-100) and suitability classifications for four procedures: Femtosecond LASIK, SMILE, TransPRK, and ICL. Five LLMs (DeepSeek-Chat, GLM-4.7, GPT-4o, Kimi-K2-Thinking, Qwen-Max) and one intermediate physician independently evaluated all cases via a structured expert-mimicking prompt. Performance was assessed using accuracy, AUC, Cohen’s kappa, correlation coefficients (R), and regression errors (RMSE, MAE). Response time and cost per query were also analyzed. LLMs demonstrated superior or comparable performance to the intermediate physician across tasks. In binary classification, top-performing LLMs achieved accuracies >98.5% and AUCs >0.96 for LASIK and SMILE. For multi-class agreement, Qwen-Max showed the highest consistency with experts (kappa up to 0.743). However, multi-class consistency (Cohen’s kappa) was more modest, indicating LLMs remain best suited as decision-support tools. In regression tasks, Qwen-Max and DeepSeek-Chat exhibited strong correlation with expert scores. The intermediate physician showed significantly lower performance, particularly in complex classifications. DeepSeek-Chat offered the best cost-efficiency with lowest cost and fastest speed, while GPT-4o was the most expensive. Advanced LLMs show strong potential as clinical decision-support tools in refractive surgery planning, with top-performing models approaching intermediate-level physician performance in binary classification tasks. However, their multi-class agreement remains moderate, and they should be positioned as assistive tools rather than autonomous decision-makers. Cost and efficiency advantages make them particularly suitable for large-scale screening and resource-limited settings.

Authors

Institutions

Publication Details

Journal
BMC Medicine
Published
2026-09-25
DOI
https://doi.org/10.1186/s12916-026-05262-4
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmarking advanced large language models for refractive surgery recommendation: a multi-model, real-world evaluation

Ran Wei, Qi Wan, Ying-ping Deng, Ke Ma et al.
BMC Medicine
Artificial Intelligence in Healthcare and Education
article

Benchmarking advanced large language models for refractive surgery recommendation: a multi-model, real-world evaluation

Ran Wei, Qi Wan, Ying-ping Deng, Ke Ma, Jing Tang
article en

Abstract

To comprehensively evaluate the performance of multiple advanced large language models (LLMs) in simulating real-world clinical decision-making for refractive surgery, comparing their accuracy, consistency, and cost-effectiveness against an intermediate-level ophthalmologist. This retrospective study analyzed preoperative data from 11,966 consecutive patients. A gold standard was established by a panel of three senior refractive surgeons (Inter-expert agreement was excellent; kappa > 0.85) who provided recommendation scores (0-100) and suitability classifications for four procedures: Femtosecond LASIK, SMILE, TransPRK, and ICL. Five LLMs (DeepSeek-Chat, GLM-4.7, GPT-4o, Kimi-K2-Thinking, Qwen-Max) and one intermediate physician independently evaluated all cases via a structured expert-mimicking prompt. Performance was assessed using accuracy, AUC, Cohen’s kappa, correlation coefficients (R), and regression errors (RMSE, MAE). Response time and cost per query were also analyzed. LLMs demonstrated superior or comparable performance to the intermediate physician across tasks. In binary classification, top-performing LLMs achieved accuracies >98.5% and AUCs >0.96 for LASIK and SMILE. For multi-class agreement, Qwen-Max showed the highest consistency with experts (kappa up to 0.743). However, multi-class consistency (Cohen’s kappa) was more modest, indicating LLMs remain best suited as decision-support tools. In regression tasks, Qwen-Max and DeepSeek-Chat exhibited strong correlation with expert scores. The intermediate physician showed significantly lower performance, particularly in complex classifications. DeepSeek-Chat offered the best cost-efficiency with lowest cost and fastest speed, while GPT-4o was the most expensive. Advanced LLMs show strong potential as clinical decision-support tools in refractive surgery planning, with top-performing models approaching intermediate-level physician performance in binary classification tasks. However, their multi-class agreement remains moderate, and they should be positioned as assistive tools rather than autonomous decision-makers. Cost and efficiency advantages make them particularly suitable for large-scale screening and resource-limited settings.

BMC Medicine
Sichuan University (CN), West China Hospital of Sichuan University (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.