Large language models reveal a systematic preference in clinical immunotherapy guidelines: a multi-model comparative analysis

Abstract Clinical guidelines may sometimes contain divergent treatment recommendations, creating uncertainty in therapeutic decision-making. To evaluate the performance of four large language models (LLMs) (GPT-4o, Claude 3.7 Sonnet, Gemini 1.5 Pro, Deepseek-V3) in analyzing discrepancies between the American Society of Clinical Oncology (ASCO) and the National Comprehensive Cancer Network (NCCN) immune checkpoint inhibitors (ICIs) guidelines. ICI-related guidelines published by ASCO and NCCN were collected, divergent recommendations were extracted and input into four LLMs for statistical analysis, and each model’s responses were independently evaluated by three oncology experts using Likert scales (focusing on accuracy, comprehensiveness, readability, and potential harm). All models demonstrated a significant preference for NCCN. Gemini 1.5 Pro exhibited extreme rigidity (100% selection of NCCN). The remaining models showed moderate degrees of a systematic preference: Claude 3.7 Sonnet(73.7%) , Deepseek-V3(72.0%), and GPT-4o(71.3%) . Further stratification revealed a relative preference for ASCO guidelines under more conservative treatment protocols. Generative artificial intelligence (AI) demonstrates persistent and systematic preference in conflicting guideline interpretation, with an overall inclination toward NCCN guidelines, potentially related to training corpus distribution and differences in guideline articulation. These findings suggest the need to establish transparent and regulatory-compliant AI decision-making frameworks.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-10-08
DOI
https://doi.org/10.1038/s41598-026-73694-2
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Large language models reveal a systematic preference in clinical immunotherapy guidelines: a multi-model comparative analysis

Shengkun Peng, Haoxuan Ying, Wenyi Gan, Qing Zeng et al.
Scientific Reports
Artificial Intelligence in Healthcare and Education
article

Large language models reveal a systematic preference in clinical immunotherapy guidelines: a multi-model comparative analysis

Shengkun Peng, Haoxuan Ying, Wenyi Gan, Qing Zeng, Zhenyu Chen, Quan Cheng, Hengguo Zhang, Hank Z. H. Wong, Mingjia Xiao, Weiming Mou, Guangdi Chu, Aimin Jiang, Junyi Shen, Wentao Xu, Chang Qi, Dongqiang Zeng, Xinpei Deng, Xuanye Cao, Bufu Tang, Xiao Liu, Xiang Wang, Wenjin Chen, Lingxuan Zhu, Lin Zhang, Peng Luo, Anqi Lin, Jian Zhang
article en

Abstract

Abstract Clinical guidelines may sometimes contain divergent treatment recommendations, creating uncertainty in therapeutic decision-making. To evaluate the performance of four large language models (LLMs) (GPT-4o, Claude 3.7 Sonnet, Gemini 1.5 Pro, Deepseek-V3) in analyzing discrepancies between the American Society of Clinical Oncology (ASCO) and the National Comprehensive Cancer Network (NCCN) immune checkpoint inhibitors (ICIs) guidelines. ICI-related guidelines published by ASCO and NCCN were collected, divergent recommendations were extracted and input into four LLMs for statistical analysis, and each model’s responses were independently evaluated by three oncology experts using Likert scales (focusing on accuracy, comprehensiveness, readability, and potential harm). All models demonstrated a significant preference for NCCN. Gemini 1.5 Pro exhibited extreme rigidity (100% selection of NCCN). The remaining models showed moderate degrees of a systematic preference: Claude 3.7 Sonnet(73.7%) , Deepseek-V3(72.0%), and GPT-4o(71.3%) . Further stratification revealed a relative preference for ASCO guidelines under more conservative treatment protocols. Generative artificial intelligence (AI) demonstrates persistent and systematic preference in conflicting guideline interpretation, with an overall inclination toward NCCN guidelines, potentially related to training corpus distribution and differences in guideline articulation. These findings suggest the need to establish transparent and regulatory-compliant AI decision-making frameworks.

Scientific Reports
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.