RadVLM: a multitask conversational vision-language model for radiology
Abstract The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-assisted reporting. While existing vision-language models (VLMs) show promise in specific tasks such as report generation or abnormality detection, they often lack support for interactive diagnostic capabilities. In this work we present RadVLM, a compact, multitask conversational VLM for CXR interpretation. We construct and standardize a large-scale CXR instruction dataset comprising over 1 million image-instruction pairs from multiple public datasets. The dataset integrates single-turn tasks—including report generation, abnormality classification, and visual grounding—with synthetic multi-turn conversations generated from structured CXR attributes. After fine-tuning RadVLM on this instruction dataset, we evaluate it across different tasks together with re-implemented baseline VLMs. Among the evaluated baselines, RadVLM achieves the strongest performance in conversational capabilities and visual grounding, while remaining competitive in other radiology tasks. Ablations comparing task-specific fine-tuning with full multitask fine-tuning are consistent with a benefit of joint training, particularly for lower-resource grounding and conversational settings. Together, these findings support RadVLM as a research prototype for structured CXR interpretation and conversational capabilities to support more effective and accessible diagnostic workflows.
Authors
- Moritz Vandenhirtz
- Hidetoshi Matsuo (ORCID: https://orcid.org/0000-0002-9684-4632)
- Michael Krauthammer (ORCID: https://orcid.org/0000-0002-4808-1845)
- Thomas Frauenfelder (ORCID: https://orcid.org/0000-0002-3295-6619)
- Thomas M. Sutter (ORCID: https://orcid.org/0000-0001-7503-4473)
- Nicolas Deperrois (ORCID: https://orcid.org/0000-0001-7178-1818)
- Sonia Laguna (ORCID: https://orcid.org/0000-0003-3504-2051)
- Julia E. Vogt (ORCID: https://orcid.org/0000-0002-6004-7770)
- Koji Fujimoto (ORCID: https://orcid.org/0000-0003-1209-7949)
- Samuel Ruipérez-Campillo (ORCID: https://orcid.org/0000-0002-5425-4175)
- Farhad Nooralahzadeh (ORCID: https://orcid.org/0000-0002-9053-0894)
- Jonas Kluckert (ORCID: https://orcid.org/0009-0008-2757-2507)
- Alain Ryser
- Christian Blüthgen (ORCID: https://orcid.org/0000-0001-7321-5676)
- Mizuho Nishio (ORCID: https://orcid.org/0000-0002-6037-2338)
Institutions
- University of Zurich (CH)
- ETH Zurich (CH)
- Kyoto College of Medical Science (JP)
- University Hospital of Zurich (CH)
- Kobe University (JP)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-08-25
- DOI
- https://doi.org/10.1038/s41598-026-66181-1
- Citations
- 2
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 6.94
Funders
- Universität Zürich
- Japan Society for the Promotion of Science