Assessing safety and trustworthiness of large language models in medicine

Abstract Background The remarkable capabilities of large language models make them increasingly compelling for use in real-world healthcare applications. However, the risks associated with using these artificial intelligence systems in medicine are not systematically understood. The aim of this study is to characterize these risks by applying five key principles for safe and trustworthy medical artificial intelligence: truthfulness, resilience, fairness, robustness, and privacy. Methods We introduce MedGuard-Bench: a safety benchmark featuring one thousand expert-verified questions covering ten specific aspects of our five core principles. We use this comprehensive corpus to systematically evaluate sixteen commonly used large language models, assessing their safety and reliability in medical contexts. Results We show that current large language models generally perform poorly on most of our safety tests, regardless of their safety alignment mechanisms. Our evaluation demonstrates that these models fall significantly short when compared to the high performance and reliability of human physicians. Conclusions Despite reports indicating that advanced large language models can match or exceed human performance in various medical tasks, this study reveals a significant safety gap in current technology. This underscores the crucial need for ongoing human oversight and the implementation of strict safety guardrails before deploying these tools in clinical practice.

Authors

Publication Details

Journal
Communications Medicine
Published
2026-09-19
DOI
https://doi.org/10.1038/s43856-026-01881-3
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Assessing safety and trustworthiness of large language models in medicine

Matías Stockle, Aidong Zhang, Changlin Gong, Furong Huang et al.
Communications Medicine
Artificial Intelligence in Healthcare and Education
article

Assessing safety and trustworthiness of large language models in medicine

Matías Stockle, Aidong Zhang, Changlin Gong, Furong Huang, Robert Leaman, Guangzhi Xiong, Jiaxin Yuan, Kelvin Castro Neira, Zhiyong Lu, Santiago Ferrière-Steinert, Maame Sarfo-Gyamfi, Xiaoyu Liu, Qiao Jin, Yifan Yang, W. John Wilbur, Bang An, Francisco Erramuspe Álvarez, Xiaojun Li
article en

Abstract

Abstract Background The remarkable capabilities of large language models make them increasingly compelling for use in real-world healthcare applications. However, the risks associated with using these artificial intelligence systems in medicine are not systematically understood. The aim of this study is to characterize these risks by applying five key principles for safe and trustworthy medical artificial intelligence: truthfulness, resilience, fairness, robustness, and privacy. Methods We introduce MedGuard-Bench: a safety benchmark featuring one thousand expert-verified questions covering ten specific aspects of our five core principles. We use this comprehensive corpus to systematically evaluate sixteen commonly used large language models, assessing their safety and reliability in medical contexts. Results We show that current large language models generally perform poorly on most of our safety tests, regardless of their safety alignment mechanisms. Our evaluation demonstrates that these models fall significantly short when compared to the high performance and reliability of human physicians. Conclusions Despite reports indicating that advanced large language models can match or exceed human performance in various medical tasks, this study reveals a significant safety gap in current technology. This underscores the crucial need for ongoing human oversight and the implementation of strict safety guardrails before deploying these tools in clinical practice.

Communications Medicine
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.