Commercial large language models for oral cavity cancer staging using descriptive pre-treatment MRI reports: ready for standalone use in clinical practice?

Abstract Objectives MRI is used for staging head and neck cancer (HNC), but assigning T- and N-category criteria requires specialised expertise, and so many institutions offer only descriptive MRI reports. This study assessed potentials of commercial large language models (LLMs) to stage oral cavity cancer (OCC) using descriptive MRI reports and compared their accuracy with human experts. Materials and methods 104 eligible MRI reports were processed by five commercial LLMs (ChatGPT5.4, ChatGPT5.0, ChatGPT4.1, Gemini3.1 and DeepSeekV3.2). T- and N-categories, and overall stage were extracted from outputs of the LLMs. MRI reports were also staged by two multidisciplinary team (MDT) members. Accuracies of the LLMs and MDT members for cancer staging were assessed against the reference staging standard and compared using the McNemar test. Results The LLMs showed accuracy of 54.8–76.0% for T-categorisation, 63.5–88.5% for N-categorisation, and 53.8–80.8% for overall stage. Compared with DeepSeekV3.2, ChatGPT and Gemini3.1 showed significantly higher accuracy ( p ≤ 0.001), except for Gemini3.1 for T-categorisation ( p = 0.07). No differences in accuracy for staging between ChatGPT versions and between them and Gemini3.1 ( p = 0.08 to > 0.99). MDT members showed accuracy of 74.0–76.9% for T-categorisation, 76.9–77.9% for N-categorisation and 68.3–71.2% for overall stage. The tested LLMs did not consistently outperform MDT members for staging. Conclusion Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. Additionally, Low MDT staging accuracy highlighted the need for structured reports that include clinical cancer staging to facilitate MDT assessment, thus potentially ensuring optimised disease management. Key Points Question Can commercial LLMs accurately assign cancer stage based on descriptive MRI reports to overcome the lack of specialised expertise in clinical practice? Findings Tested LLMs demonstrated variable performances, ranging 54.8–76.0% for T-categorisation, 63.5–88.5% for N-categorisation and 53.8–80.8% for overall stage, and failed to consistently outperform MDT members. Critical relevance statement Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. MDT members’ low performances highlight that structured MRI reports with clinical cancer staging are needed to improve multidisciplinary assessment.

Authors

Institutions

Publication Details

Journal
Insights into Imaging
Published
2026-09-21
DOI
https://doi.org/10.1186/s13244-026-02409-y
Primary Topic
Head and Neck Cancer Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Commercial large language models for oral cavity cancer staging using descriptive pre-treatment MRI reports: ready for standalone use in clinical practice?

Qi Yong H. Ai, Hoi Ming Kwok, Kuo Feng Hung, Tiffany Y. So et al.
Insights into Imaging
Head and Neck Cancer Studies
article

Commercial large language models for oral cavity cancer staging using descriptive pre-treatment MRI reports: ready for standalone use in clinical practice?

Qi Yong H. Ai, Hoi Ming Kwok, Kuo Feng Hung, Tiffany Y. So, Ho Sang Leung, Ann D. King, Lun M. Wong, Tracy T. S. Lau, Ming-Yi Lu
article en

Abstract

Abstract Objectives MRI is used for staging head and neck cancer (HNC), but assigning T- and N-category criteria requires specialised expertise, and so many institutions offer only descriptive MRI reports. This study assessed potentials of commercial large language models (LLMs) to stage oral cavity cancer (OCC) using descriptive MRI reports and compared their accuracy with human experts. Materials and methods 104 eligible MRI reports were processed by five commercial LLMs (ChatGPT5.4, ChatGPT5.0, ChatGPT4.1, Gemini3.1 and DeepSeekV3.2). T- and N-categories, and overall stage were extracted from outputs of the LLMs. MRI reports were also staged by two multidisciplinary team (MDT) members. Accuracies of the LLMs and MDT members for cancer staging were assessed against the reference staging standard and compared using the McNemar test. Results The LLMs showed accuracy of 54.8–76.0% for T-categorisation, 63.5–88.5% for N-categorisation, and 53.8–80.8% for overall stage. Compared with DeepSeekV3.2, ChatGPT and Gemini3.1 showed significantly higher accuracy ( p ≤ 0.001), except for Gemini3.1 for T-categorisation ( p = 0.07). No differences in accuracy for staging between ChatGPT versions and between them and Gemini3.1 ( p = 0.08 to > 0.99). MDT members showed accuracy of 74.0–76.9% for T-categorisation, 76.9–77.9% for N-categorisation and 68.3–71.2% for overall stage. The tested LLMs did not consistently outperform MDT members for staging. Conclusion Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. Additionally, Low MDT staging accuracy highlighted the need for structured reports that include clinical cancer staging to facilitate MDT assessment, thus potentially ensuring optimised disease management. Key Points Question Can commercial LLMs accurately assign cancer stage based on descriptive MRI reports to overcome the lack of specialised expertise in clinical practice? Findings Tested LLMs demonstrated variable performances, ranging 54.8–76.0% for T-categorisation, 63.5–88.5% for N-categorisation and 53.8–80.8% for overall stage, and failed to consistently outperform MDT members. Critical relevance statement Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. MDT members’ low performances highlight that structured MRI reports with clinical cancer staging are needed to improve multidisciplinary assessment.

Insights into ImagingVol. 17(1)
Hospital Authority (HK), Chinese University of Hong Kong (HK), Chung Shan Medical University Hospital (TW), Prince of Wales Hospital (CN), Princess Margaret Hospital (NZ), University of Hong Kong (HK), Chung Shan Medical University (TW)
Openalex Percentile: Top 8%
Head and Neck Cancer Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.