Large Language Model Benchmarks in Medical Tasks

ABSTRACT With the increasing application of large language models (LLMs) in the medical domain, evaluating these models' performance using benchmark datasets has become crucial. This paper presents a comprehensive survey of various benchmark datasets used in medical LLM tasks. These datasets span multiple modalities including text, image, and multimodal benchmarks, focusing on various aspects of medical knowledge such as electronic health records, doctor–patient dialogues, medical question answering, and medical image captioning. The survey categorizes the datasets by modality and examines their significance, data structure, and roles in model development and evaluation across tasks such as diagnostic support, report generation, and predictive decision support. Representative resources include Medical Information Mart for Intensive Care III (MIMIC‐III), MIMIC‐IV, BioASQ, PubMedQA, and CheXpert, which provide data and evaluation settings for research in clinical NLP, medical question answering, and chest‐radiograph interpretation. This paper summarizes the challenges and opportunities in leveraging these benchmarks for advancing multimodal medical intelligence, emphasizing the need for datasets with a greater degree of language diversity, structured omics data, and innovative approaches to synthesis. This synthesis is intended to inform future research on the applications of LLMs in medicine and medical artificial intelligence.

Authors

Institutions

Publication Details

Journal
Medicine Advances
Published
2026-09-19
DOI
https://doi.org/10.1002/med4.70085
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Large Language Model Benchmarks in Medical Tasks

Riyang Bao, Zekun Jiang, Benji Peng, Lawrence K.Q. Yan et al.
Medicine Advances
Artificial Intelligence in Healthcare and Education
article

Large Language Model Benchmarks in Medical Tasks

Riyang Bao, Zekun Jiang, Benji Peng, Lawrence K.Q. Yan, Cheng Fei, Silin Chen, Junyu Liu, Ziyuan Qin, Ziqian Bi, Yichao Zhang, Xinyuan Song, Keyu Chen, Qian Niu, Ming Li, Yunze Wang, Tianyang Wang, Ming Liu, Pohsun Feng, Caitlyn Heqi Yin
article en

Abstract

ABSTRACT With the increasing application of large language models (LLMs) in the medical domain, evaluating these models' performance using benchmark datasets has become crucial. This paper presents a comprehensive survey of various benchmark datasets used in medical LLM tasks. These datasets span multiple modalities including text, image, and multimodal benchmarks, focusing on various aspects of medical knowledge such as electronic health records, doctor–patient dialogues, medical question answering, and medical image captioning. The survey categorizes the datasets by modality and examines their significance, data structure, and roles in model development and evaluation across tasks such as diagnostic support, report generation, and predictive decision support. Representative resources include Medical Information Mart for Intensive Care III (MIMIC‐III), MIMIC‐IV, BioASQ, PubMedQA, and CheXpert, which provide data and evaluation settings for research in clinical NLP, medical question answering, and chest‐radiograph interpretation. This paper summarizes the challenges and opportunities in leveraging these benchmarks for advancing multimodal medical intelligence, emphasizing the need for datasets with a greater degree of language diversity, structured omics data, and innovative approaches to synthesis. This synthesis is intended to inform future research on the applications of LLMs in medicine and medical artificial intelligence.

Medicine Advances
Georgia Institute of Technology (US), National Taiwan Normal University (TW), University of Wisconsin–Madison (US), University of Liverpool (GB), Emory University (US), The University of Texas at Dallas (US), Hong Kong University of Science and Technology (HK), Cornell University (US), Purdue University West Lafayette (US), Kyoto University (JP), Sichuan University (CN), West China Hospital of Sichuan University (CN), Acupuncture And Massage College (US), Second Affiliated Hospital of Zhejiang University (CN), University of Edinburgh (GB)
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.