Performance of 4 large language models across different disciplines in the Chinese National Medical Licensing Examination

Large language models (LLMs) have shown strong performance on medical licensing examinations, but differences across disciplines and inter-model agreement remain insufficiently characterized. We evaluated 1,635 Chinese National Medical Licensing Examination multiple-choice questions, including 129 basic medicine, 838 clinical medicine, and 668 medical humanities items, using four model-platform configurations: DeepSeek-V3.2, ChatGPT-5.2, Gemini 3 Pro, and Claude 4.5 Sonnet. Accuracy was assessed across modules, subdisciplines, question types, and difficulty levels, with paired item-level analyses for medical humanities overall, health law, and A2 questions in clinical medicine and medical humanities. Mixed-effects logistic regression and inter-model agreement analyses were also performed. Overall accuracy was high across all four configurations. Significant differences were found in medical humanities overall and health law, as well as for A2 questions in clinical medicine and medical humanities. Accuracy declined with increasing item difficulty, and A2 questions were associated with lower odds of a correct response than A1 questions. All four configurations selected the same final answer for 79.9% of items, with an overall Fleiss’ κ of 0.856. These findings indicate high overall performance but also differences in specific subgroups, lower accuracy on more difficult questions, and incomplete item-level agreement among model-platform configurations.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-15
DOI
https://doi.org/10.1038/s41598-026-71907-2
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Performance of 4 large language models across different disciplines in the Chinese National Medical Licensing Examination

Yuning Zhang, Lingxia Chen, Qi Xu, Xiaolu Xie
Scientific Reports
Artificial Intelligence in Healthcare and Education
article

Performance of 4 large language models across different disciplines in the Chinese National Medical Licensing Examination

Yuning Zhang, Lingxia Chen, Qi Xu, Xiaolu Xie
article en

Abstract

Large language models (LLMs) have shown strong performance on medical licensing examinations, but differences across disciplines and inter-model agreement remain insufficiently characterized. We evaluated 1,635 Chinese National Medical Licensing Examination multiple-choice questions, including 129 basic medicine, 838 clinical medicine, and 668 medical humanities items, using four model-platform configurations: DeepSeek-V3.2, ChatGPT-5.2, Gemini 3 Pro, and Claude 4.5 Sonnet. Accuracy was assessed across modules, subdisciplines, question types, and difficulty levels, with paired item-level analyses for medical humanities overall, health law, and A2 questions in clinical medicine and medical humanities. Mixed-effects logistic regression and inter-model agreement analyses were also performed. Overall accuracy was high across all four configurations. Significant differences were found in medical humanities overall and health law, as well as for A2 questions in clinical medicine and medical humanities. Accuracy declined with increasing item difficulty, and A2 questions were associated with lower odds of a correct response than A1 questions. All four configurations selected the same final answer for 79.9% of items, with an overall Fleiss’ κ of 0.856. These findings indicate high overall performance but also differences in specific subgroups, lower accuracy on more difficult questions, and incomplete item-level agreement among model-platform configurations.

Scientific Reports
Gannan Medical University (CN)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.