Performance of large language models in the Chinese National Medical Licensing Examinations and Beyond: Scoping review
Large language models (LLMs) such as GPT-4 and DeepSeek have shown potential in medical education and assessment. In China, the Chinese National Medical Licensing Examination (CNMLE) has become a critical benchmark for evaluating LLMs’ professional knowledge and reasoning ability in non-English contexts. This scoping review aims to synthesize research evaluating LLM performance on the CNMLE and related Chinese medical examinations, identifying performance trends, methodological gaps, and future directions. A scoping review was conducted in accordance with PRISMA-ScR guidelines. PubMed and Web of Science were searched on June 12, 2025, using terms related to LLMs and Chinese medical examinations. Studies were included if they evaluated any LLM on national-level Chinese medical exams and reported performance metrics. Two reviewers screened and extracted data. Study quality was assessed using a 12-item checklist covering dataset characteristics, LLM setup, and evaluation methods. Fisher’s exact test was used to assess differences in the passing rates of different LLMs. 14 studies were included, covering 51 evaluation records across 8 types of Chinese medical examinations, including CNMLE and specialty exams such as critical care, radiation oncology, and ultrasound medicine. Exam years ranged from 2017 to 2024, with a shift toward using recent exams. A total of 9 LLMs were evaluated, including GPT-3.5, GPT-4, GPT-4o, ERNIE, DeepSeek-R1, Qwen-72B, Baichuan2-7B, Baichuan2-13B and DISC-MedLLM. GPT-4o and DeepSeek-R1 achieved the highest scores of 552 (92.00%) and 523 (87.20%) in the 2024 CNMLE, respectively, both surpassing the passing threshold. In an exploratory comparison, GPT-4 achieved a higher pass rate than GPT-3.5 with a statistically significant difference, although inconsistencies in datasets, model transparency, and evaluation criteria persist across the included studies. LLMs show improving performance on Chinese medical exams, especially newer models. Nevertheless, future research should prioritize standardized benchmarks, multimodal capabilities, and transparent evaluation to ensure meaningful clinical relevance and educational value.
Authors
- Jiaxue Cha (ORCID: https://orcid.org/0000-0002-5917-5441)
- Hui Zong (ORCID: https://orcid.org/0000-0002-9142-5017)
- Bairong Shen (ORCID: https://orcid.org/0000-0003-2899-1531)
- Rongrong Wu
- Shanshan Hu
- Yan Zhao
- Jiao Wang
Institutions
- Tongji University (CN)
- Army Medical University (CN)
- Southwest Hospital (CN)
- First Affiliated Hospital of Soochow University (CN)
- Artificial Intelligence in Medicine (Canada) (CA)
Publication Details
- Journal
- PLOS Digital Health
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1371/journal.pdig.0001773
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00