Benchmarking and developing large language models using one million clinical trials
Abstract Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation for model benchmarking and development. Here, we introduce , a large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature. Using this resource, we construct 152K training and testing samples spanning eight clinical research tasks, including systematic review, trial design, and trial optimization. Benchmarking cutting-edge large language models (LLMs) reveals limited clinical reasoning capability in generic LLMs. In contrast, an 8B LLM developed on using supervised fine-tuning and reinforcement learning outperforms 70B generic counterparts across all eight tasks, with relative improvements of 73.7, 67.6, 38.4, 37.8, 26.5, 20.7, 20.0, 18.1, and 5.2%, respectively. These results demonstrate the potential of domain-adapted AI to improve evidence synthesis and clinical trial design, establishing as a foundation for scaling AI in clinical research.
Authors
- Jimeng Sun (ORCID: https://orcid.org/0000-0003-1512-6426)
- Jiacheng Lin
- Qiao Jin
- Zhiyong Lu
- Zifeng Wang
- Jathurshan Pradeepkumar
- Pengcheng Jiang
- Junyi Gao
Institutions
- National Institutes of Health (US)
- University of Illinois Urbana-Champaign (US)
- United States National Library of Medicine (US)
- Keiju Medical Center (JP)
- Health Data Research UK (GB)
- University of Edinburgh (GB)
Publication Details
- Journal
- npj Digital Medicine
- Published
- 2026-07-31
- DOI
- https://doi.org/10.1038/s41746-026-02933-7
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00