RadVLM: a multitask conversational vision-language model for radiology

Abstract The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-assisted reporting. While existing vision-language models (VLMs) show promise in specific tasks such as report generation or abnormality detection, they often lack support for interactive diagnostic capabilities. In this work we present RadVLM, a compact, multitask conversational VLM for CXR interpretation. We construct and standardize a large-scale CXR instruction dataset comprising over 1 million image-instruction pairs from multiple public datasets. The dataset integrates single-turn tasks—including report generation, abnormality classification, and visual grounding—with synthetic multi-turn conversations generated from structured CXR attributes. After fine-tuning RadVLM on this instruction dataset, we evaluate it across different tasks together with re-implemented baseline VLMs. Among the evaluated baselines, RadVLM achieves the strongest performance in conversational capabilities and visual grounding, while remaining competitive in other radiology tasks. Ablations comparing task-specific fine-tuning with full multitask fine-tuning are consistent with a benefit of joint training, particularly for lower-resource grounding and conversational settings. Together, these findings support RadVLM as a research prototype for structured CXR interpretation and conversational capabilities to support more effective and accessible diagnostic workflows.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-08-25
DOI
https://doi.org/10.1038/s41598-026-66181-1
Citations
2
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
6.94

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

RadVLM: a multitask conversational vision-language model for radiology

Moritz Vandenhirtz, Hidetoshi Matsuo, Michael Krauthammer, Thomas Frauenfelder et al.
2 citations
Scientific Reports
Topic Modeling
6.94
article

RadVLM: a multitask conversational vision-language model for radiology

Moritz Vandenhirtz, Hidetoshi Matsuo, Michael Krauthammer, Thomas Frauenfelder, Thomas M. Sutter, Nicolas Deperrois, Sonia Laguna, Julia E. Vogt, Koji Fujimoto, Samuel Ruipérez-Campillo, Farhad Nooralahzadeh, Jonas Kluckert, Alain Ryser, Christian Blüthgen, Mizuho Nishio
article en
2 citations

Abstract

Abstract The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-assisted reporting. While existing vision-language models (VLMs) show promise in specific tasks such as report generation or abnormality detection, they often lack support for interactive diagnostic capabilities. In this work we present RadVLM, a compact, multitask conversational VLM for CXR interpretation. We construct and standardize a large-scale CXR instruction dataset comprising over 1 million image-instruction pairs from multiple public datasets. The dataset integrates single-turn tasks—including report generation, abnormality classification, and visual grounding—with synthetic multi-turn conversations generated from structured CXR attributes. After fine-tuning RadVLM on this instruction dataset, we evaluate it across different tasks together with re-implemented baseline VLMs. Among the evaluated baselines, RadVLM achieves the strongest performance in conversational capabilities and visual grounding, while remaining competitive in other radiology tasks. Ablations comparing task-specific fine-tuning with full multitask fine-tuning are consistent with a benefit of joint training, particularly for lower-resource grounding and conversational settings. Together, these findings support RadVLM as a research prototype for structured CXR interpretation and conversational capabilities to support more effective and accessible diagnostic workflows.

Scientific Reports
University of Zurich (CH), ETH Zurich (CH), Kyoto College of Medical Science (JP), University Hospital of Zurich (CH), Kobe University (JP)
Universität Zürich, Japan Society for the Promotion of Science
Openalex Percentile: Top 7%
Topic Modeling
6.94
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.