Promise Versus Performance: A Faithful Benchmark of Vision-Language Foundation Models Against Specialized AI for Lung Cancer Prediction
Vision-language foundation models (VLMs) have demonstrated broad competence across medical imaging tasks, raising the question of whether they can match purpose-built, specialised AI systems for high-stakes 3D CT screening tasks. This talk presents a retrospective benchmark comparing eight VLM configurations — MedGemma, M3FM, CT-CHAT, MedSigLIP, and COLIPRI (each evaluated zero-shot and/or fine-tuned) — against seven specialized, CT-native cancer detection AI models on the NLST held-out test set (n=2,065; 54 cancers). Specialized models consistently exceed AUC 0.94, led by the Eyonis LCS ensemble at 0.98, while VLMs range from near-chance zero-shot performance to AUC 0.84-0.91 depending on fine-tuning and prompt design — notably, COLIPRI's malignancy-grounded zero-shot prompting reaches AUC ~0.91 at a fraction of the compute cost (~15 TFLOPs/scan) of larger VLMs (up to 6,700 TFLOPs/scan). The key driver of this gap is CT-native domain pretraining, not model scale or general vision-language capability: COLIPRI, the only VLM in our benchmark pretrained natively on CT volumes, is also the only one to approach specialised-level AUC, while Pillar-0's CT-native vision-only encoder matches top CNN specialists using nothing more than a linear probe on frozen embeddings. To ground these results in clinical relevance, I will also draw on two companion studies benchmarking against radiologists directly. In a 12-reader study, native MedGemma (AUC 0.70) fell short of clinically relevant performance, while fine-tuned MedGemma (AUC 0.83) reached only the level of less-experienced readers (mean radiologist AUC 0.90; range 0.80-0.94). By contrast, in an independent multi-reader evaluation, the specialised Eyonis LCS AI outperformed all radiologists. Combined with the markedly higher computational burden of generic vision-language models, this gap currently favours CT-native specialised systems for safer near-term clinical use, with fewer diagnostic errors and unnecessary procedures. For the SAFER community, this underscores that faithful evaluation of foundation models must be paired with genuine domain adaptation — or, better still, dedicated 3D CT pretraining from the outset — not just scale, before deployment in screening pipelines.
Authors
- Ezequiel Geremia
- Benoît Huet (ORCID: https://orcid.org/0000-0002-0608-6939)
- Pierre Baudot (ORCID: https://orcid.org/0000-0002-5574-6809)
- Benjamin Renoust (ORCID: https://orcid.org/0000-0003-2692-9056)
- Jean‐Christophe Brisset (ORCID: https://orcid.org/0000-0002-7947-3622)
- Danny Francis (ORCID: https://orcid.org/0009-0009-3185-0693)
Institutions
- Media Health Technologies (United States) (US)
- Median (Czechia) (CZ)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-24
- DOI
- https://doi.org/10.5281/zenodo.22942251
- Primary Topic
- Lung Cancer Diagnosis and Treatment
- Type
- article
- Field-Weighted Citation Impact
- 0.00