An open vision-language model for diverse medical applications

Abstract Artificial intelligence has high potential for impact in healthcare applications, but its training and deployment are challenging due to diverse data, a complex spectrum of possible tasks and important privacy needs. High-performing foundation models that enable data-efficient fine-tuning for diverse downstream tasks can meaningfully accelerate development in this domain. We introduce MedGemma, a collection of medical vision-language foundation models based on Gemma 3. MedGemma demonstrates advanced medical understanding and reasoning across images and text and multiple medical imaging domains, exceeding the performance of similarly sized generative models while maintaining the general capabilities of the Gemma base models. For out-of-distribution tasks, MedGemma achieves improvements of 2.6–10% in medical image question answering, 15.5–18.1% in chest X-ray finding classification and 10.8% in agentic evaluations compared with the base models. Our results show that fine-tuning MedGemma can be more effective than fine-tuning the base Gemma 3 model for medical tasks, particularly in the setting of limited training data. We additionally introduce MedSigLIP, a medically tuned vision encoder derived from SigLIP. MedSigLIP powers the visual understanding capabilities of MedGemma and, as an encoder, achieves performance comparable to or better than that of many specialized medical image encoders. Taken together, the MedGemma collection provides a strong foundation of medical image and text capabilities, with the potential to accelerate medical research and the development of downstream applications.

Authors

Institutions

Publication Details

Journal
Nature Medicine
Published
2026-10-06
DOI
https://doi.org/10.1038/s41591-026-04626-w
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

An open vision-language model for diverse medical applications

Shravya Shetty, David F. Steiner, Yubin Kim, Daniel I Golden et al.
Nature Medicine
Multimodal Machine Learning Applications
article

An open vision-language model for diverse medical applications

Shravya Shetty, David F. Steiner, Yubin Kim, Daniel I Golden, Jean-Baptiste Alayrac, Rachelle Sico, Cían Owen Hughes, Atilla P. Kiraly, Clément Farabet, Geoffrey Cideron, Jonathon Shlens, Elena Buchatskaya, David J. Fleet, Sabela Ramos, Andrew Sellergren, Can Kirmizibayrak, Robert Dadashi, Joelle K. Barral, Shekoofeh Azizi, Dmitry Lepikhin, Kenneth A. Philbrick, Charles T. Lau, Tiffany Chen, Jean-Bastien Grill, Louis Rouillard, Alexandre Ramé, Timo Kohlberger, Sunny Jansen, Sebastian Borgeaud, Dale R. Webster, Sathaiah Baby, Yossi Matias, Cassidy Hardin, Liron Yatziv, Katherine Chou, Mercy Asiedu, Nino Vieillard, Sahar Kazemzadeh, Justin Anthony Chen, Samuel Schmidgall, Aishwarya Kamath, Yun Liu, Léonard Hussenot, Bram Sterling, Edouard Yvinec, Shawn Xu, Rory Pilgrim, Johan Ferret, Omar Sanseviero, Olivier Bachem, Ronnachai Jaroensri, Vlad Feinberg, Tatiana Matejovicova, Tris Warkentin, Victor Cotruta, Daniel McDuff, Jeremy Lai, Richa Tiwari, Avinatan Hassidim, Chufan Gao, Ramona Merhej, Michelle Casbon, Armand Joulin, Shashir Reddy, Morgane Rivière, Gus Martins, Alek Andreev, Phoebe Kirk, Fereshteh Mahvar, Thomas Mesnard, Ryan Brush, Per Bjornsson, Ines Mezerreg, Madeleine Traverse, Kejia Chen, Dr. Anand Rao, Susanna Maria Baby, Shreya Pathak, Kavi Goel, Howard Yang, Fayaz Jamil, Howard Hu, Catherine Kozlowski, Sarah Perrin, Lu Yang, Preeti Singh, Lin Yang, Bashir Sadjad
article en

Abstract

Abstract Artificial intelligence has high potential for impact in healthcare applications, but its training and deployment are challenging due to diverse data, a complex spectrum of possible tasks and important privacy needs. High-performing foundation models that enable data-efficient fine-tuning for diverse downstream tasks can meaningfully accelerate development in this domain. We introduce MedGemma, a collection of medical vision-language foundation models based on Gemma 3. MedGemma demonstrates advanced medical understanding and reasoning across images and text and multiple medical imaging domains, exceeding the performance of similarly sized generative models while maintaining the general capabilities of the Gemma base models. For out-of-distribution tasks, MedGemma achieves improvements of 2.6–10% in medical image question answering, 15.5–18.1% in chest X-ray finding classification and 10.8% in agentic evaluations compared with the base models. Our results show that fine-tuning MedGemma can be more effective than fine-tuning the base Gemma 3 model for medical tasks, particularly in the setting of limited training data. We additionally introduce MedSigLIP, a medically tuned vision encoder derived from SigLIP. MedSigLIP powers the visual understanding capabilities of MedGemma and, as an encoder, achieves performance comparable to or better than that of many specialized medical image encoders. Taken together, the MedGemma collection provides a strong foundation of medical image and text capabilities, with the potential to accelerate medical research and the development of downstream applications.

Nature Medicine
Google (United States) (US)
Openalex Percentile: Top 15%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.