Accuracy of large language models in head and neck cancers: a comparative analysis of ChatGPT and Gemini in TNM staging and clinical decision support

Objective Large Language Models (LLMs) are increasingly integrated into oncological workflows. This study evaluated and compared the performance of ChatGPT-4o and Gemini 1.5 Pro in TNM staging and treatment planning for head and neck malignancies across diverse anatomical sites. Materials and Methods A retrospective analysis was performed on 180 patients with head and neck cancer. Clinical data were processed through structured prompts to simulate real-world clinical inquiries. AI-generated TNM stages and treatment protocols were compared against AJCC 8th Edition guidelines and expert multidisciplinary tumor board decisions. The analysis focused on diagnostic accuracy, inter-rater agreement, and the impact of anatomical complexity on model performance. Results Both models demonstrated comparable proficiency in TNM staging accuracy (ChatGPT: 75.6%, Gemini: 75.0%), showing substantial agreement with expert standards (x 2 = 0.000, p = 1.000). However, a significant divergence was observed in treatment planning; Gemini achieved a 78.9% accuracy rate, significantly outperforming ChatGPT’s 71.7% (p=0.043 (x 2 = 4.114). Notably, ChatGPT’s staging performance was sensitive to tumor localization, with decreased precision in anatomically complex regions such as the oropharynx and paranasal sinuses (p = 0.034, Cramer’s V = 0.291). Conversely, Gemini demonstrated more robust spatial reasoning across different subsites. Conclusion While both LLMs provide reliable staging support, Gemini exhibits superior clinical reasoning in synthesizing multidimensional data into actionable treatment recommendations. However, a staging error rate of 25% remains a critical concern, potentially leading to inappropriate clinical pathways. These models should be viewed as auxiliary tools within an ‘augmented intelligence’ ecosystem, integrated with imaging and multidisciplinary inputs, rather than independent decision-makers. Strict expert supervision is mandatory to prevent subsite-specific errors from impacting surgical and oncological outcomes.

Authors

Institutions

Publication Details

Journal
Frontiers in Oncology
Published
2026-05-08
DOI
https://doi.org/10.3389/fonc.2026.1828538
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Accuracy of large language models in head and neck cancers: a comparative analysis of ChatGPT and Gemini in TNM staging and clinical decision support

Huseyin Isik, Duygu Erdem, Gökhan F. Kılıç, Deniz Baklaci
Frontiers in Oncology
Artificial Intelligence in Healthcare and Education
article

Accuracy of large language models in head and neck cancers: a comparative analysis of ChatGPT and Gemini in TNM staging and clinical decision support

Huseyin Isik, Duygu Erdem, Gökhan F. Kılıç, Deniz Baklaci
article en

Abstract

Objective Large Language Models (LLMs) are increasingly integrated into oncological workflows. This study evaluated and compared the performance of ChatGPT-4o and Gemini 1.5 Pro in TNM staging and treatment planning for head and neck malignancies across diverse anatomical sites. Materials and Methods A retrospective analysis was performed on 180 patients with head and neck cancer. Clinical data were processed through structured prompts to simulate real-world clinical inquiries. AI-generated TNM stages and treatment protocols were compared against AJCC 8th Edition guidelines and expert multidisciplinary tumor board decisions. The analysis focused on diagnostic accuracy, inter-rater agreement, and the impact of anatomical complexity on model performance. Results Both models demonstrated comparable proficiency in TNM staging accuracy (ChatGPT: 75.6%, Gemini: 75.0%), showing substantial agreement with expert standards (x 2 = 0.000, p = 1.000). However, a significant divergence was observed in treatment planning; Gemini achieved a 78.9% accuracy rate, significantly outperforming ChatGPT’s 71.7% (p=0.043 (x 2 = 4.114). Notably, ChatGPT’s staging performance was sensitive to tumor localization, with decreased precision in anatomically complex regions such as the oropharynx and paranasal sinuses (p = 0.034, Cramer’s V = 0.291). Conversely, Gemini demonstrated more robust spatial reasoning across different subsites. Conclusion While both LLMs provide reliable staging support, Gemini exhibits superior clinical reasoning in synthesizing multidimensional data into actionable treatment recommendations. However, a staging error rate of 25% remains a critical concern, potentially leading to inappropriate clinical pathways. These models should be viewed as auxiliary tools within an ‘augmented intelligence’ ecosystem, integrated with imaging and multidisciplinary inputs, rather than independent decision-makers. Strict expert supervision is mandatory to prevent subsite-specific errors from impacting surgical and oncological outcomes.

Frontiers in OncologyVol. 16
Zonguldak Bülent Ecevit University (TR)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.