Multimodal Large Language Models for Prognostic Prediction in Cervical Cancer Treated with Definitive Chemoradiotherapy: An Exploratory Study of Systematic Multimodal Data Integration
Background/Objectives: About 25–35% of patients with cervical cancer treated with definitive concurrent chemoradiotherapy (CCRT) relapse within five years, and non-imaging clinical factors stratify their risk only modestly. We evaluated whether a general-purpose multimodal large language model (MLLM), used without fine-tuning, could estimate recurrence risk in this setting. Methods: In this retrospective single-center study, 82 patients treated with definitive radiotherapy (79/82 with concurrent platinum-based chemotherapy) who had a complete pretreatment MRI report were analyzed. Gemini 3.1 Pro generated a structured report from pelvic MRI images and, separately, estimated recurrence/metastasis risk from clinical, laboratory, treatment and imaging data. Because treatment cycles and overall treatment time actually completed were included, this was a retrospective treatment-complete risk assessment rather than a strictly pretreatment prediction. Five prompting strategies differing only in their inputs—an ablation of input modalities—were each run three times; AI reports were graded against paired human reports on a 14-item rubric by an independent model (Claude Opus 4.6). Results: Forty patients (48.8%) relapsed. Discrimination rose from an AUC of 0.704 with a clinical baseline to 0.790 with the full multimodal input (95% CI 0.687–0.883; ΔAUC +0.085). This gain was significant on our primary two-sided bootstrap test after Holm correction for ten pairwise comparisons (p = 0.032; a DeLong sensitivity analysis supported some but not all the imaging-benefit comparisons). Adding MRI report text significantly improved the AUC over the clinical baseline, and no statistically significant difference was detected between the human-report and AI-report strategies; this was not an equivalence or non-inferiority test; and the further lymph-node increment was not statistically significant. The AI reports scored 68.8% against an LLM judge on the rubric, with apparent overcalling of parametrial invasion and an inability to assess lymph nodes within the supplied field of view; all strategies underestimated absolute risk (E/O 0.78–0.83). Conclusions: A non-fine-tuned MLLM can integrate multimodal data into a prognostic estimate, but its automatically generated MRI reports are frequently discordant with radiologist reports in specific ways and require expert review and external validation before clinical use.
Authors
- Yidong Zhang (ORCID: https://orcid.org/0000-0002-6966-9306)
- Zhaoqi Gu (ORCID: https://orcid.org/0000-0003-4225-7229)
- Qizhen Zhu
- Weiping Wang
- Chen Wang
- Ke Hu
Institutions
- Chinese Academy of Medical Sciences & Peking Union Medical College (CN)
- Peking Union Medical College Hospital (CN)
Publication Details
- Journal
- Cancers
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/cancers18193136
- Primary Topic
- Endometrial and Cervical Cancer Treatments
- Type
- article
- Field-Weighted Citation Impact
- 0.00