Promise Versus Performance: A Faithful Benchmark of Vision-Language Foundation Models Against Specialized AI for Lung Cancer Prediction

Vision-language foundation models (VLMs) have demonstrated broad competence across medical imaging tasks, raising the question of whether they can match purpose-built, specialised AI systems for high-stakes 3D CT screening tasks. This talk presents a retrospective benchmark comparing eight VLM configurations — MedGemma, M3FM, CT-CHAT, MedSigLIP, and COLIPRI (each evaluated zero-shot and/or fine-tuned) — against seven specialized, CT-native cancer detection AI models on the NLST held-out test set (n=2,065; 54 cancers). Specialized models consistently exceed AUC 0.94, led by the Eyonis LCS ensemble at 0.98, while VLMs range from near-chance zero-shot performance to AUC 0.84-0.91 depending on fine-tuning and prompt design — notably, COLIPRI's malignancy-grounded zero-shot prompting reaches AUC ~0.91 at a fraction of the compute cost (~15 TFLOPs/scan) of larger VLMs (up to 6,700 TFLOPs/scan). The key driver of this gap is CT-native domain pretraining, not model scale or general vision-language capability: COLIPRI, the only VLM in our benchmark pretrained natively on CT volumes, is also the only one to approach specialised-level AUC, while Pillar-0's CT-native vision-only encoder matches top CNN specialists using nothing more than a linear probe on frozen embeddings. To ground these results in clinical relevance, I will also draw on two companion studies benchmarking against radiologists directly. In a 12-reader study, native MedGemma (AUC 0.70) fell short of clinically relevant performance, while fine-tuned MedGemma (AUC 0.83) reached only the level of less-experienced readers (mean radiologist AUC 0.90; range 0.80-0.94). By contrast, in an independent multi-reader evaluation, the specialised Eyonis LCS AI outperformed all radiologists. Combined with the markedly higher computational burden of generic vision-language models, this gap currently favours CT-native specialised systems for safer near-term clinical use, with fewer diagnostic errors and unnecessary procedures. For the SAFER community, this underscores that faithful evaluation of foundation models must be paired with genuine domain adaptation — or, better still, dedicated 3D CT pretraining from the outset — not just scale, before deployment in screening pipelines.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-24
DOI
https://doi.org/10.5281/zenodo.22942251
Primary Topic
Lung Cancer Diagnosis and Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Promise Versus Performance: A Faithful Benchmark of Vision-Language Foundation Models Against Specialized AI for Lung Cancer Prediction

Ezequiel Geremia, Benoît Huet, Pierre Baudot, Benjamin Renoust et al.
Zenodo (CERN European Organization for Nuclear Research)
Lung Cancer Diagnosis and Treatment
article

Promise Versus Performance: A Faithful Benchmark of Vision-Language Foundation Models Against Specialized AI for Lung Cancer Prediction

Ezequiel Geremia, Benoît Huet, Pierre Baudot, Benjamin Renoust, Jean‐Christophe Brisset, Danny Francis
article en

Abstract

Vision-language foundation models (VLMs) have demonstrated broad competence across medical imaging tasks, raising the question of whether they can match purpose-built, specialised AI systems for high-stakes 3D CT screening tasks. This talk presents a retrospective benchmark comparing eight VLM configurations — MedGemma, M3FM, CT-CHAT, MedSigLIP, and COLIPRI (each evaluated zero-shot and/or fine-tuned) — against seven specialized, CT-native cancer detection AI models on the NLST held-out test set (n=2,065; 54 cancers). Specialized models consistently exceed AUC 0.94, led by the Eyonis LCS ensemble at 0.98, while VLMs range from near-chance zero-shot performance to AUC 0.84-0.91 depending on fine-tuning and prompt design — notably, COLIPRI's malignancy-grounded zero-shot prompting reaches AUC ~0.91 at a fraction of the compute cost (~15 TFLOPs/scan) of larger VLMs (up to 6,700 TFLOPs/scan). The key driver of this gap is CT-native domain pretraining, not model scale or general vision-language capability: COLIPRI, the only VLM in our benchmark pretrained natively on CT volumes, is also the only one to approach specialised-level AUC, while Pillar-0's CT-native vision-only encoder matches top CNN specialists using nothing more than a linear probe on frozen embeddings. To ground these results in clinical relevance, I will also draw on two companion studies benchmarking against radiologists directly. In a 12-reader study, native MedGemma (AUC 0.70) fell short of clinically relevant performance, while fine-tuned MedGemma (AUC 0.83) reached only the level of less-experienced readers (mean radiologist AUC 0.90; range 0.80-0.94). By contrast, in an independent multi-reader evaluation, the specialised Eyonis LCS AI outperformed all radiologists. Combined with the markedly higher computational burden of generic vision-language models, this gap currently favours CT-native specialised systems for safer near-term clinical use, with fewer diagnostic errors and unnecessary procedures. For the SAFER community, this underscores that faithful evaluation of foundation models must be paired with genuine domain adaptation — or, better still, dedicated 3D CT pretraining from the outset — not just scale, before deployment in screening pipelines.

Zenodo (CERN European Organization for Nuclear Research)
Media Health Technologies (United States) (US), Median (Czechia) (CZ)
Quality Education
Openalex Percentile: Top 12%
Lung Cancer Diagnosis and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.