Assessing safety and trustworthiness of large language models in medicine
Abstract Background The remarkable capabilities of large language models make them increasingly compelling for use in real-world healthcare applications. However, the risks associated with using these artificial intelligence systems in medicine are not systematically understood. The aim of this study is to characterize these risks by applying five key principles for safe and trustworthy medical artificial intelligence: truthfulness, resilience, fairness, robustness, and privacy. Methods We introduce MedGuard-Bench: a safety benchmark featuring one thousand expert-verified questions covering ten specific aspects of our five core principles. We use this comprehensive corpus to systematically evaluate sixteen commonly used large language models, assessing their safety and reliability in medical contexts. Results We show that current large language models generally perform poorly on most of our safety tests, regardless of their safety alignment mechanisms. Our evaluation demonstrates that these models fall significantly short when compared to the high performance and reliability of human physicians. Conclusions Despite reports indicating that advanced large language models can match or exceed human performance in various medical tasks, this study reveals a significant safety gap in current technology. This underscores the crucial need for ongoing human oversight and the implementation of strict safety guardrails before deploying these tools in clinical practice.
Authors
- Matías Stockle
- Aidong Zhang (ORCID: https://orcid.org/0000-0001-9723-3246)
- Changlin Gong (ORCID: https://orcid.org/0000-0001-9932-0538)
- Furong Huang (ORCID: https://orcid.org/0000-0002-7393-1047)
- Robert Leaman (ORCID: https://orcid.org/0000-0003-3296-5766)
- Guangzhi Xiong (ORCID: https://orcid.org/0000-0002-8049-5298)
- Jiaxin Yuan (ORCID: https://orcid.org/0000-0003-1739-2381)
- Kelvin Castro Neira (ORCID: https://orcid.org/0000-0002-2871-8183)
- Zhiyong Lu (ORCID: https://orcid.org/0000-0001-9998-916X)
- Santiago Ferrière-Steinert (ORCID: https://orcid.org/0009-0009-8296-4076)
- Maame Sarfo-Gyamfi
- Xiaoyu Liu (ORCID: https://orcid.org/0000-0002-3485-0885)
- Qiao Jin (ORCID: https://orcid.org/0000-0002-1268-7239)
- Yifan Yang (ORCID: https://orcid.org/0000-0003-4414-9176)
- W. John Wilbur
- Bang An
- Francisco Erramuspe Álvarez
- Xiaojun Li
Publication Details
- Journal
- Communications Medicine
- Published
- 2026-09-19
- DOI
- https://doi.org/10.1038/s43856-026-01881-3
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00