Commercial large language models for oral cavity cancer staging using descriptive pre-treatment MRI reports: ready for standalone use in clinical practice?
Abstract Objectives MRI is used for staging head and neck cancer (HNC), but assigning T- and N-category criteria requires specialised expertise, and so many institutions offer only descriptive MRI reports. This study assessed potentials of commercial large language models (LLMs) to stage oral cavity cancer (OCC) using descriptive MRI reports and compared their accuracy with human experts. Materials and methods 104 eligible MRI reports were processed by five commercial LLMs (ChatGPT5.4, ChatGPT5.0, ChatGPT4.1, Gemini3.1 and DeepSeekV3.2). T- and N-categories, and overall stage were extracted from outputs of the LLMs. MRI reports were also staged by two multidisciplinary team (MDT) members. Accuracies of the LLMs and MDT members for cancer staging were assessed against the reference staging standard and compared using the McNemar test. Results The LLMs showed accuracy of 54.8–76.0% for T-categorisation, 63.5–88.5% for N-categorisation, and 53.8–80.8% for overall stage. Compared with DeepSeekV3.2, ChatGPT and Gemini3.1 showed significantly higher accuracy ( p ≤ 0.001), except for Gemini3.1 for T-categorisation ( p = 0.07). No differences in accuracy for staging between ChatGPT versions and between them and Gemini3.1 ( p = 0.08 to > 0.99). MDT members showed accuracy of 74.0–76.9% for T-categorisation, 76.9–77.9% for N-categorisation and 68.3–71.2% for overall stage. The tested LLMs did not consistently outperform MDT members for staging. Conclusion Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. Additionally, Low MDT staging accuracy highlighted the need for structured reports that include clinical cancer staging to facilitate MDT assessment, thus potentially ensuring optimised disease management. Key Points Question Can commercial LLMs accurately assign cancer stage based on descriptive MRI reports to overcome the lack of specialised expertise in clinical practice? Findings Tested LLMs demonstrated variable performances, ranging 54.8–76.0% for T-categorisation, 63.5–88.5% for N-categorisation and 53.8–80.8% for overall stage, and failed to consistently outperform MDT members. Critical relevance statement Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. MDT members’ low performances highlight that structured MRI reports with clinical cancer staging are needed to improve multidisciplinary assessment.
Authors
- Qi Yong H. Ai (ORCID: https://orcid.org/0000-0002-1155-4389)
- Hoi Ming Kwok (ORCID: https://orcid.org/0000-0001-7628-1687)
- Kuo Feng Hung (ORCID: https://orcid.org/0000-0002-3971-3484)
- Tiffany Y. So (ORCID: https://orcid.org/0000-0001-8268-0721)
- Ho Sang Leung
- Ann D. King
- Lun M. Wong
- Tracy T. S. Lau
- Ming-Yi Lu
Institutions
- Hospital Authority (HK)
- Chinese University of Hong Kong (HK)
- Chung Shan Medical University Hospital (TW)
- Prince of Wales Hospital (CN)
- Princess Margaret Hospital (NZ)
- University of Hong Kong (HK)
- Chung Shan Medical University (TW)
Publication Details
- Journal
- Insights into Imaging
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1186/s13244-026-02409-y
- Primary Topic
- Head and Neck Cancer Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00