A multicenter assessment of human oversight of generative AI outputs in simulated clinical decision making

Generative AI, particularly Large Language Models (LLMs), is increasingly being explored for selected clinical tasks, including clinical documentation, patient communication, and decision support; however, routine use in direct clinical decision-making remains limited. A central barrier to safe deployment, however, is the problem of medical hallucinations. Current safeguards rely on a “clinician-in- the-loop” model, which assumes that clinicians can reliably identify and correct these hallucinations. Whether this assumption actually holds in clinical practice warrants greater empirical attention, particularly among junior clinicians. In this multi-center cross-sectional study, we systematically detected and classified GPT-4o-generated hallucinations across diverse clinical scenarios. We also quantitatively evaluated junior clinicians’ actual identification capabilities across varying clinical categories and risk levels. Our results revealed that only 15.8% of hallucinations were identified, and 13.1% of clinicians failed to detect any hallucinations across all scenarios. Notably, detection rates did not improve with increasing clinical risk, and variability in hallucination detection arose primarily from between-clinician rather than between-scenario differences. Our findings reveal a critical gap between technical validation and real-world clinical use. These results indicate that future clinical applications will require well-designed, structured human-AI collaborative workflows, along with tiered clinical certification pathways, to ensure patient safety.

Authors

Institutions

Publication Details

Journal
npj Digital Medicine
Published
2026-09-24
DOI
https://doi.org/10.1038/s41746-026-03294-x
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A multicenter assessment of human oversight of generative AI outputs in simulated clinical decision making

Bo Yue, Yunyun Zhang, Jintao Zhang, Xiaochuan Cui et al.
npj Digital Medicine
Artificial Intelligence in Healthcare and Education
article

A multicenter assessment of human oversight of generative AI outputs in simulated clinical decision making

Bo Yue, Yunyun Zhang, Jintao Zhang, Xiaochuan Cui, Rongrong Wan, Bingbing Fu, Jiacheng Zhou, Zhiyong Zhang, Jia Meng, Hua Guo
article en

Abstract

Generative AI, particularly Large Language Models (LLMs), is increasingly being explored for selected clinical tasks, including clinical documentation, patient communication, and decision support; however, routine use in direct clinical decision-making remains limited. A central barrier to safe deployment, however, is the problem of medical hallucinations. Current safeguards rely on a “clinician-in- the-loop” model, which assumes that clinicians can reliably identify and correct these hallucinations. Whether this assumption actually holds in clinical practice warrants greater empirical attention, particularly among junior clinicians. In this multi-center cross-sectional study, we systematically detected and classified GPT-4o-generated hallucinations across diverse clinical scenarios. We also quantitatively evaluated junior clinicians’ actual identification capabilities across varying clinical categories and risk levels. Our results revealed that only 15.8% of hallucinations were identified, and 13.1% of clinicians failed to detect any hallucinations across all scenarios. Notably, detection rates did not improve with increasing clinical risk, and variability in hallucination detection arose primarily from between-clinician rather than between-scenario differences. Our findings reveal a critical gap between technical validation and real-world clinical use. These results indicate that future clinical applications will require well-designed, structured human-AI collaborative workflows, along with tiered clinical certification pathways, to ensure patient safety.

npj Digital Medicine
Harbin Medical University (CN), First Affiliated Hospital of Jiamusi University (CN), Second Affiliated Hospital of Harbin Medical University (CN), Wuxi People's Hospital (CN), Qiqihar Medical University (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.