Performance of large language models in the Chinese National Medical Licensing Examinations and Beyond: Scoping review

Large language models (LLMs) such as GPT-4 and DeepSeek have shown potential in medical education and assessment. In China, the Chinese National Medical Licensing Examination (CNMLE) has become a critical benchmark for evaluating LLMs’ professional knowledge and reasoning ability in non-English contexts. This scoping review aims to synthesize research evaluating LLM performance on the CNMLE and related Chinese medical examinations, identifying performance trends, methodological gaps, and future directions. A scoping review was conducted in accordance with PRISMA-ScR guidelines. PubMed and Web of Science were searched on June 12, 2025, using terms related to LLMs and Chinese medical examinations. Studies were included if they evaluated any LLM on national-level Chinese medical exams and reported performance metrics. Two reviewers screened and extracted data. Study quality was assessed using a 12-item checklist covering dataset characteristics, LLM setup, and evaluation methods. Fisher’s exact test was used to assess differences in the passing rates of different LLMs. 14 studies were included, covering 51 evaluation records across 8 types of Chinese medical examinations, including CNMLE and specialty exams such as critical care, radiation oncology, and ultrasound medicine. Exam years ranged from 2017 to 2024, with a shift toward using recent exams. A total of 9 LLMs were evaluated, including GPT-3.5, GPT-4, GPT-4o, ERNIE, DeepSeek-R1, Qwen-72B, Baichuan2-7B, Baichuan2-13B and DISC-MedLLM. GPT-4o and DeepSeek-R1 achieved the highest scores of 552 (92.00%) and 523 (87.20%) in the 2024 CNMLE, respectively, both surpassing the passing threshold. In an exploratory comparison, GPT-4 achieved a higher pass rate than GPT-3.5 with a statistically significant difference, although inconsistencies in datasets, model transparency, and evaluation criteria persist across the included studies. LLMs show improving performance on Chinese medical exams, especially newer models. Nevertheless, future research should prioritize standardized benchmarks, multimodal capabilities, and transparent evaluation to ensure meaningful clinical relevance and educational value.

Authors

Institutions

Publication Details

Journal
PLOS Digital Health
Published
2026-10-09
DOI
https://doi.org/10.1371/journal.pdig.0001773
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Performance of large language models in the Chinese National Medical Licensing Examinations and Beyond: Scoping review

Jiaxue Cha, Hui Zong, Bairong Shen, Rongrong Wu et al.
PLOS Digital Health
Artificial Intelligence in Healthcare and Education
article

Performance of large language models in the Chinese National Medical Licensing Examinations and Beyond: Scoping review

Jiaxue Cha, Hui Zong, Bairong Shen, Rongrong Wu, Shanshan Hu, Yan Zhao, Jiao Wang
article en

Abstract

Large language models (LLMs) such as GPT-4 and DeepSeek have shown potential in medical education and assessment. In China, the Chinese National Medical Licensing Examination (CNMLE) has become a critical benchmark for evaluating LLMs’ professional knowledge and reasoning ability in non-English contexts. This scoping review aims to synthesize research evaluating LLM performance on the CNMLE and related Chinese medical examinations, identifying performance trends, methodological gaps, and future directions. A scoping review was conducted in accordance with PRISMA-ScR guidelines. PubMed and Web of Science were searched on June 12, 2025, using terms related to LLMs and Chinese medical examinations. Studies were included if they evaluated any LLM on national-level Chinese medical exams and reported performance metrics. Two reviewers screened and extracted data. Study quality was assessed using a 12-item checklist covering dataset characteristics, LLM setup, and evaluation methods. Fisher’s exact test was used to assess differences in the passing rates of different LLMs. 14 studies were included, covering 51 evaluation records across 8 types of Chinese medical examinations, including CNMLE and specialty exams such as critical care, radiation oncology, and ultrasound medicine. Exam years ranged from 2017 to 2024, with a shift toward using recent exams. A total of 9 LLMs were evaluated, including GPT-3.5, GPT-4, GPT-4o, ERNIE, DeepSeek-R1, Qwen-72B, Baichuan2-7B, Baichuan2-13B and DISC-MedLLM. GPT-4o and DeepSeek-R1 achieved the highest scores of 552 (92.00%) and 523 (87.20%) in the 2024 CNMLE, respectively, both surpassing the passing threshold. In an exploratory comparison, GPT-4 achieved a higher pass rate than GPT-3.5 with a statistically significant difference, although inconsistencies in datasets, model transparency, and evaluation criteria persist across the included studies. LLMs show improving performance on Chinese medical exams, especially newer models. Nevertheless, future research should prioritize standardized benchmarks, multimodal capabilities, and transparent evaluation to ensure meaningful clinical relevance and educational value.

PLOS Digital HealthVol. 5(10)
Tongji University (CN), Army Medical University (CN), Southwest Hospital (CN), First Affiliated Hospital of Soochow University (CN), Artificial Intelligence in Medicine (Canada) (CA)
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.