Large language models reveal a systematic preference in clinical immunotherapy guidelines: a multi-model comparative analysis
Abstract Clinical guidelines may sometimes contain divergent treatment recommendations, creating uncertainty in therapeutic decision-making. To evaluate the performance of four large language models (LLMs) (GPT-4o, Claude 3.7 Sonnet, Gemini 1.5 Pro, Deepseek-V3) in analyzing discrepancies between the American Society of Clinical Oncology (ASCO) and the National Comprehensive Cancer Network (NCCN) immune checkpoint inhibitors (ICIs) guidelines. ICI-related guidelines published by ASCO and NCCN were collected, divergent recommendations were extracted and input into four LLMs for statistical analysis, and each model’s responses were independently evaluated by three oncology experts using Likert scales (focusing on accuracy, comprehensiveness, readability, and potential harm). All models demonstrated a significant preference for NCCN. Gemini 1.5 Pro exhibited extreme rigidity (100% selection of NCCN). The remaining models showed moderate degrees of a systematic preference: Claude 3.7 Sonnet(73.7%) , Deepseek-V3(72.0%), and GPT-4o(71.3%) . Further stratification revealed a relative preference for ASCO guidelines under more conservative treatment protocols. Generative artificial intelligence (AI) demonstrates persistent and systematic preference in conflicting guideline interpretation, with an overall inclination toward NCCN guidelines, potentially related to training corpus distribution and differences in guideline articulation. These findings suggest the need to establish transparent and regulatory-compliant AI decision-making frameworks.
Authors
- Shengkun Peng
- Haoxuan Ying (ORCID: https://orcid.org/0000-0002-0957-9827)
- Wenyi Gan (ORCID: https://orcid.org/0000-0003-1886-8062)
- Qing Zeng (ORCID: https://orcid.org/0000-0002-7471-4941)
- Zhenyu Chen (ORCID: https://orcid.org/0000-0002-3642-3385)
- Quan Cheng (ORCID: https://orcid.org/0000-0003-2401-5349)
- Hengguo Zhang (ORCID: https://orcid.org/0000-0002-4438-8348)
- Hank Z. H. Wong
- Mingjia Xiao
- Weiming Mou (ORCID: https://orcid.org/0009-0007-1089-6516)
- Guangdi Chu (ORCID: https://orcid.org/0000-0002-2963-1741)
- Aimin Jiang (ORCID: https://orcid.org/0000-0002-9563-983X)
- Junyi Shen (ORCID: https://orcid.org/0000-0001-8990-867X)
- Wentao Xu (ORCID: https://orcid.org/0000-0003-0679-1213)
- Chang Qi
- Dongqiang Zeng
- Xinpei Deng
- Xuanye Cao
- Bufu Tang
- Xiao Liu
- Xiang Wang
- Wenjin Chen
- Lingxuan Zhu
- Lin Zhang
- Peng Luo
- Anqi Lin
- Jian Zhang
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1038/s41598-026-73694-2
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00