Multimodal Large Language Models for Prognostic Prediction in Cervical Cancer Treated with Definitive Chemoradiotherapy: An Exploratory Study of Systematic Multimodal Data Integration

Background/Objectives: About 25–35% of patients with cervical cancer treated with definitive concurrent chemoradiotherapy (CCRT) relapse within five years, and non-imaging clinical factors stratify their risk only modestly. We evaluated whether a general-purpose multimodal large language model (MLLM), used without fine-tuning, could estimate recurrence risk in this setting. Methods: In this retrospective single-center study, 82 patients treated with definitive radiotherapy (79/82 with concurrent platinum-based chemotherapy) who had a complete pretreatment MRI report were analyzed. Gemini 3.1 Pro generated a structured report from pelvic MRI images and, separately, estimated recurrence/metastasis risk from clinical, laboratory, treatment and imaging data. Because treatment cycles and overall treatment time actually completed were included, this was a retrospective treatment-complete risk assessment rather than a strictly pretreatment prediction. Five prompting strategies differing only in their inputs—an ablation of input modalities—were each run three times; AI reports were graded against paired human reports on a 14-item rubric by an independent model (Claude Opus 4.6). Results: Forty patients (48.8%) relapsed. Discrimination rose from an AUC of 0.704 with a clinical baseline to 0.790 with the full multimodal input (95% CI 0.687–0.883; ΔAUC +0.085). This gain was significant on our primary two-sided bootstrap test after Holm correction for ten pairwise comparisons (p = 0.032; a DeLong sensitivity analysis supported some but not all the imaging-benefit comparisons). Adding MRI report text significantly improved the AUC over the clinical baseline, and no statistically significant difference was detected between the human-report and AI-report strategies; this was not an equivalence or non-inferiority test; and the further lymph-node increment was not statistically significant. The AI reports scored 68.8% against an LLM judge on the rubric, with apparent overcalling of parametrial invasion and an inability to assess lymph nodes within the supplied field of view; all strategies underestimated absolute risk (E/O 0.78–0.83). Conclusions: A non-fine-tuned MLLM can integrate multimodal data into a prognostic estimate, but its automatically generated MRI reports are frequently discordant with radiologist reports in specific ways and require expert review and external validation before clinical use.

Authors

Institutions

Publication Details

Journal
Cancers
Published
2026-09-28
DOI
https://doi.org/10.3390/cancers18193136
Primary Topic
Endometrial and Cervical Cancer Treatments
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multimodal Large Language Models for Prognostic Prediction in Cervical Cancer Treated with Definitive Chemoradiotherapy: An Exploratory Study of Systematic Multimodal Data Integration

Yidong Zhang, Zhaoqi Gu, Qizhen Zhu, Weiping Wang et al.
Cancers
Endometrial and Cervical Cancer Treatments
article

Multimodal Large Language Models for Prognostic Prediction in Cervical Cancer Treated with Definitive Chemoradiotherapy: An Exploratory Study of Systematic Multimodal Data Integration

Yidong Zhang, Zhaoqi Gu, Qizhen Zhu, Weiping Wang, Chen Wang, Ke Hu
article en

Abstract

Background/Objectives: About 25–35% of patients with cervical cancer treated with definitive concurrent chemoradiotherapy (CCRT) relapse within five years, and non-imaging clinical factors stratify their risk only modestly. We evaluated whether a general-purpose multimodal large language model (MLLM), used without fine-tuning, could estimate recurrence risk in this setting. Methods: In this retrospective single-center study, 82 patients treated with definitive radiotherapy (79/82 with concurrent platinum-based chemotherapy) who had a complete pretreatment MRI report were analyzed. Gemini 3.1 Pro generated a structured report from pelvic MRI images and, separately, estimated recurrence/metastasis risk from clinical, laboratory, treatment and imaging data. Because treatment cycles and overall treatment time actually completed were included, this was a retrospective treatment-complete risk assessment rather than a strictly pretreatment prediction. Five prompting strategies differing only in their inputs—an ablation of input modalities—were each run three times; AI reports were graded against paired human reports on a 14-item rubric by an independent model (Claude Opus 4.6). Results: Forty patients (48.8%) relapsed. Discrimination rose from an AUC of 0.704 with a clinical baseline to 0.790 with the full multimodal input (95% CI 0.687–0.883; ΔAUC +0.085). This gain was significant on our primary two-sided bootstrap test after Holm correction for ten pairwise comparisons (p = 0.032; a DeLong sensitivity analysis supported some but not all the imaging-benefit comparisons). Adding MRI report text significantly improved the AUC over the clinical baseline, and no statistically significant difference was detected between the human-report and AI-report strategies; this was not an equivalence or non-inferiority test; and the further lymph-node increment was not statistically significant. The AI reports scored 68.8% against an LLM judge on the rubric, with apparent overcalling of parametrial invasion and an inability to assess lymph nodes within the supplied field of view; all strategies underestimated absolute risk (E/O 0.78–0.83). Conclusions: A non-fine-tuned MLLM can integrate multimodal data into a prognostic estimate, but its automatically generated MRI reports are frequently discordant with radiologist reports in specific ways and require expert review and external validation before clinical use.

CancersVol. 18(19)
Chinese Academy of Medical Sciences & Peking Union Medical College (CN), Peking Union Medical College Hospital (CN)
Reduced inequalities
Openalex Percentile: Top 8%
Endometrial and Cervical Cancer Treatments
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.