A multicenter assessment of human oversight of generative AI outputs in simulated clinical decision making
Generative AI, particularly Large Language Models (LLMs), is increasingly being explored for selected clinical tasks, including clinical documentation, patient communication, and decision support; however, routine use in direct clinical decision-making remains limited. A central barrier to safe deployment, however, is the problem of medical hallucinations. Current safeguards rely on a “clinician-in- the-loop” model, which assumes that clinicians can reliably identify and correct these hallucinations. Whether this assumption actually holds in clinical practice warrants greater empirical attention, particularly among junior clinicians. In this multi-center cross-sectional study, we systematically detected and classified GPT-4o-generated hallucinations across diverse clinical scenarios. We also quantitatively evaluated junior clinicians’ actual identification capabilities across varying clinical categories and risk levels. Our results revealed that only 15.8% of hallucinations were identified, and 13.1% of clinicians failed to detect any hallucinations across all scenarios. Notably, detection rates did not improve with increasing clinical risk, and variability in hallucination detection arose primarily from between-clinician rather than between-scenario differences. Our findings reveal a critical gap between technical validation and real-world clinical use. These results indicate that future clinical applications will require well-designed, structured human-AI collaborative workflows, along with tiered clinical certification pathways, to ensure patient safety.
Authors
- Bo Yue
- Yunyun Zhang (ORCID: https://orcid.org/0000-0001-7744-0141)
- Jintao Zhang (ORCID: https://orcid.org/0000-0002-2909-4552)
- Xiaochuan Cui (ORCID: https://orcid.org/0000-0001-8903-1396)
- Rongrong Wan (ORCID: https://orcid.org/0000-0002-4705-8846)
- Bingbing Fu (ORCID: https://orcid.org/0000-0002-9122-0708)
- Jiacheng Zhou (ORCID: https://orcid.org/0000-0003-1092-3848)
- Zhiyong Zhang (ORCID: https://orcid.org/0000-0001-8202-9672)
- Jia Meng
- Hua Guo
Institutions
- Harbin Medical University (CN)
- First Affiliated Hospital of Jiamusi University (CN)
- Second Affiliated Hospital of Harbin Medical University (CN)
- Wuxi People's Hospital (CN)
- Qiqihar Medical University (CN)
Publication Details
- Journal
- npj Digital Medicine
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1038/s41746-026-03294-x
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00