Benchmarking and developing large language models using one million clinical trials

Abstract Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation for model benchmarking and development. Here, we introduce , a large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature. Using this resource, we construct 152K training and testing samples spanning eight clinical research tasks, including systematic review, trial design, and trial optimization. Benchmarking cutting-edge large language models (LLMs) reveals limited clinical reasoning capability in generic LLMs. In contrast, an 8B LLM developed on using supervised fine-tuning and reinforcement learning outperforms 70B generic counterparts across all eight tasks, with relative improvements of 73.7, 67.6, 38.4, 37.8, 26.5, 20.7, 20.0, 18.1, and 5.2%, respectively. These results demonstrate the potential of domain-adapted AI to improve evidence synthesis and clinical trial design, establishing as a foundation for scaling AI in clinical research.

Authors

Institutions

Publication Details

Journal
npj Digital Medicine
Published
2026-07-31
DOI
https://doi.org/10.1038/s41746-026-02933-7
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmarking and developing large language models using one million clinical trials

Jimeng Sun, Jiacheng Lin, Qiao Jin, Zhiyong Lu et al.
npj Digital Medicine
Artificial Intelligence in Healthcare and Education
article

Benchmarking and developing large language models using one million clinical trials

Jimeng Sun, Jiacheng Lin, Qiao Jin, Zhiyong Lu, Zifeng Wang, Jathurshan Pradeepkumar, Pengcheng Jiang, Junyi Gao
article en

Abstract

Abstract Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation for model benchmarking and development. Here, we introduce , a large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature. Using this resource, we construct 152K training and testing samples spanning eight clinical research tasks, including systematic review, trial design, and trial optimization. Benchmarking cutting-edge large language models (LLMs) reveals limited clinical reasoning capability in generic LLMs. In contrast, an 8B LLM developed on using supervised fine-tuning and reinforcement learning outperforms 70B generic counterparts across all eight tasks, with relative improvements of 73.7, 67.6, 38.4, 37.8, 26.5, 20.7, 20.0, 18.1, and 5.2%, respectively. These results demonstrate the potential of domain-adapted AI to improve evidence synthesis and clinical trial design, establishing as a foundation for scaling AI in clinical research.

npj Digital Medicine
National Institutes of Health (US), University of Illinois Urbana-Champaign (US), United States National Library of Medicine (US), Keiju Medical Center (JP), Health Data Research UK (GB), University of Edinburgh (GB)
Industry, innovation and infrastructure
Openalex Percentile: Top 12%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.