Multimodal large language models for bladder tumor detection in cystoscopy: a retrospective benchmarking study

PURPOSE: Cystoscopic assessment is central to bladder cancer diagnosis, yet visual interpretation remains variable. Existing artificial intelligence approaches often depend on data-intensive models that are often difficult to deploy in routine practice. We evaluated whether multimodal large language models (MLLMs), including smaller and more efficient architectures, can accurately classify cystoscopy images, and whether prompt engineering improves performance. The primary outcome was benign-versus-malignant classification. Secondary outcomes included calibration, high-confidence triage, and performance stratified by imaging modality. METHODS: We retrospectively analyzed 1,754 labeled public cystoscopy images. Three prompt types were tested: Direct, Book-based, and Optimized, across GPT-5.2, GPT-5, GPT-5-Mini, and GPT-5-Nano. Performance measured: accuracy, sensitivity, specificity, and F1 Score. Confidence evaluation: using Brier Score and Expected Calibration Error. High-confidence triage using abstention option based on loss function. RESULTS: GPT-5 and GPT-5-Mini with the optimized prompt achieved the best benign-versus-malignant performance, with accuracies of 86.7% and 89.2%, specificities of 94.1% and 88.4%, and sensitivities of 82.5% and 89.2%, respectively. GPT-5 with the optimized prompt achieved the best high-confidence triage performance, yielding 98.1% accuracy, 94.6% specificity, and 99.1% sensitivity at 62.0% image coverage. Prompt engineering improved model performance, although these gains were not statistically significant, and enhanced confidence calibration and triage performance. CONCLUSIONS: This retrospective evaluation demonstrates the potential of MLLMs for cystoscopic bladder lesion classification. Prompt engineering improved diagnostic calibration and output reliability, while high-confidence triage increased accuracy to 98.1%, supporting the feasibility of MLLMs as foundation models for cystoscopic assessment.

Authors

Institutions

Publication Details

Journal
World Journal of Urology
Published
2026-09-17
DOI
https://doi.org/10.1007/s00345-026-06761-y
Primary Topic
Bladder and Urothelial Cancer Treatments
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multimodal large language models for bladder tumor detection in cystoscopy: a retrospective benchmarking study

Barak Rosenzweig, Abraham Tsur, ZOHAR A. DOTAN, Menachem Laufer et al.
World Journal of Urology
Bladder and Urothelial Cancer Treatments
article

Multimodal large language models for bladder tumor detection in cystoscopy: a retrospective benchmarking study

Barak Rosenzweig, Abraham Tsur, ZOHAR A. DOTAN, Menachem Laufer, Husny Mahmud, Yonatan Prat, Dina Orkin
article en

Abstract

PURPOSE: Cystoscopic assessment is central to bladder cancer diagnosis, yet visual interpretation remains variable. Existing artificial intelligence approaches often depend on data-intensive models that are often difficult to deploy in routine practice. We evaluated whether multimodal large language models (MLLMs), including smaller and more efficient architectures, can accurately classify cystoscopy images, and whether prompt engineering improves performance. The primary outcome was benign-versus-malignant classification. Secondary outcomes included calibration, high-confidence triage, and performance stratified by imaging modality. METHODS: We retrospectively analyzed 1,754 labeled public cystoscopy images. Three prompt types were tested: Direct, Book-based, and Optimized, across GPT-5.2, GPT-5, GPT-5-Mini, and GPT-5-Nano. Performance measured: accuracy, sensitivity, specificity, and F1 Score. Confidence evaluation: using Brier Score and Expected Calibration Error. High-confidence triage using abstention option based on loss function. RESULTS: GPT-5 and GPT-5-Mini with the optimized prompt achieved the best benign-versus-malignant performance, with accuracies of 86.7% and 89.2%, specificities of 94.1% and 88.4%, and sensitivities of 82.5% and 89.2%, respectively. GPT-5 with the optimized prompt achieved the best high-confidence triage performance, yielding 98.1% accuracy, 94.6% specificity, and 99.1% sensitivity at 62.0% image coverage. Prompt engineering improved model performance, although these gains were not statistically significant, and enhanced confidence calibration and triage performance. CONCLUSIONS: This retrospective evaluation demonstrates the potential of MLLMs for cystoscopic bladder lesion classification. Prompt engineering improved diagnostic calibration and output reliability, while high-confidence triage increased accuracy to 98.1%, supporting the feasibility of MLLMs as foundation models for cystoscopic assessment.

World Journal of UrologyVol. 44(1)
Reichman University (IL), Tel Aviv University (IL), Tel Aviv Sourasky Medical Center (IL), Sheba Medical Center (IL), Herzliya Medical Center (IL)
Tel Aviv University
Openalex Percentile: Top 9%
Bladder and Urothelial Cancer Treatments
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.