Large Language Model Benchmarks in Medical Tasks
ABSTRACT With the increasing application of large language models (LLMs) in the medical domain, evaluating these models' performance using benchmark datasets has become crucial. This paper presents a comprehensive survey of various benchmark datasets used in medical LLM tasks. These datasets span multiple modalities including text, image, and multimodal benchmarks, focusing on various aspects of medical knowledge such as electronic health records, doctor–patient dialogues, medical question answering, and medical image captioning. The survey categorizes the datasets by modality and examines their significance, data structure, and roles in model development and evaluation across tasks such as diagnostic support, report generation, and predictive decision support. Representative resources include Medical Information Mart for Intensive Care III (MIMIC‐III), MIMIC‐IV, BioASQ, PubMedQA, and CheXpert, which provide data and evaluation settings for research in clinical NLP, medical question answering, and chest‐radiograph interpretation. This paper summarizes the challenges and opportunities in leveraging these benchmarks for advancing multimodal medical intelligence, emphasizing the need for datasets with a greater degree of language diversity, structured omics data, and innovative approaches to synthesis. This synthesis is intended to inform future research on the applications of LLMs in medicine and medical artificial intelligence.
Authors
- Riyang Bao (ORCID: https://orcid.org/0000-0002-3763-4539)
- Zekun Jiang (ORCID: https://orcid.org/0000-0002-3178-7761)
- Benji Peng (ORCID: https://orcid.org/0009-0003-6552-3064)
- Lawrence K.Q. Yan (ORCID: https://orcid.org/0000-0003-3400-9356)
- Cheng Fei (ORCID: https://orcid.org/0009-0008-4922-9259)
- Silin Chen
- Junyu Liu
- Ziyuan Qin
- Ziqian Bi
- Yichao Zhang
- Xinyuan Song
- Keyu Chen
- Qian Niu
- Ming Li
- Yunze Wang
- Tianyang Wang
- Ming Liu
- Pohsun Feng
- Caitlyn Heqi Yin
Institutions
- Georgia Institute of Technology (US)
- National Taiwan Normal University (TW)
- University of Wisconsin–Madison (US)
- University of Liverpool (GB)
- Emory University (US)
- The University of Texas at Dallas (US)
- Hong Kong University of Science and Technology (HK)
- Cornell University (US)
- Purdue University West Lafayette (US)
- Kyoto University (JP)
- Sichuan University (CN)
- West China Hospital of Sichuan University (CN)
- Acupuncture And Massage College (US)
- Second Affiliated Hospital of Zhejiang University (CN)
- University of Edinburgh (GB)
Publication Details
- Journal
- Medicine Advances
- Published
- 2026-09-19
- DOI
- https://doi.org/10.1002/med4.70085
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00