An open vision-language model for diverse medical applications
Abstract Artificial intelligence has high potential for impact in healthcare applications, but its training and deployment are challenging due to diverse data, a complex spectrum of possible tasks and important privacy needs. High-performing foundation models that enable data-efficient fine-tuning for diverse downstream tasks can meaningfully accelerate development in this domain. We introduce MedGemma, a collection of medical vision-language foundation models based on Gemma 3. MedGemma demonstrates advanced medical understanding and reasoning across images and text and multiple medical imaging domains, exceeding the performance of similarly sized generative models while maintaining the general capabilities of the Gemma base models. For out-of-distribution tasks, MedGemma achieves improvements of 2.6–10% in medical image question answering, 15.5–18.1% in chest X-ray finding classification and 10.8% in agentic evaluations compared with the base models. Our results show that fine-tuning MedGemma can be more effective than fine-tuning the base Gemma 3 model for medical tasks, particularly in the setting of limited training data. We additionally introduce MedSigLIP, a medically tuned vision encoder derived from SigLIP. MedSigLIP powers the visual understanding capabilities of MedGemma and, as an encoder, achieves performance comparable to or better than that of many specialized medical image encoders. Taken together, the MedGemma collection provides a strong foundation of medical image and text capabilities, with the potential to accelerate medical research and the development of downstream applications.
Authors
- Shravya Shetty (ORCID: https://orcid.org/0000-0003-3783-3172)
- David F. Steiner (ORCID: https://orcid.org/0000-0003-1297-0023)
- Yubin Kim (ORCID: https://orcid.org/0000-0002-1902-3822)
- Daniel I Golden (ORCID: https://orcid.org/0000-0002-1114-3184)
- Jean-Baptiste Alayrac (ORCID: https://orcid.org/0000-0002-3071-4157)
- Rachelle Sico
- Cían Owen Hughes (ORCID: https://orcid.org/0000-0001-6901-0985)
- Atilla P. Kiraly (ORCID: https://orcid.org/0000-0002-6613-3581)
- Clément Farabet
- Geoffrey Cideron
- Jonathon Shlens (ORCID: https://orcid.org/0000-0001-9513-4244)
- Elena Buchatskaya
- David J. Fleet (ORCID: https://orcid.org/0000-0003-0734-7114)
- Sabela Ramos (ORCID: https://orcid.org/0000-0001-6656-9732)
- Andrew Sellergren (ORCID: https://orcid.org/0000-0002-5260-8013)
- Can Kirmizibayrak
- Robert Dadashi
- Joelle K. Barral (ORCID: https://orcid.org/0009-0009-0432-5148)
- Shekoofeh Azizi (ORCID: https://orcid.org/0000-0002-7447-6031)
- Dmitry Lepikhin
- Kenneth A. Philbrick (ORCID: https://orcid.org/0000-0003-1424-6987)
- Charles T. Lau (ORCID: https://orcid.org/0000-0002-3136-9711)
- Tiffany Chen (ORCID: https://orcid.org/0000-0003-2054-2758)
- Jean-Bastien Grill
- Louis Rouillard
- Alexandre Ramé
- Timo Kohlberger (ORCID: https://orcid.org/0009-0003-7145-6128)
- Sunny Jansen (ORCID: https://orcid.org/0009-0003-4957-4486)
- Sebastian Borgeaud
- Dale R. Webster (ORCID: https://orcid.org/0000-0002-3023-8824)
- Sathaiah Baby
- Yossi Matias (ORCID: https://orcid.org/0000-0003-3960-6002)
- Cassidy Hardin
- Liron Yatziv
- Katherine Chou (ORCID: https://orcid.org/0000-0002-0318-7857)
- Mercy Asiedu (ORCID: https://orcid.org/0000-0002-0230-5022)
- Nino Vieillard
- Sahar Kazemzadeh (ORCID: https://orcid.org/0009-0008-6640-6288)
- Justin Anthony Chen (ORCID: https://orcid.org/0000-0002-0218-1867)
- Samuel Schmidgall (ORCID: https://orcid.org/0000-0001-8192-9337)
- Aishwarya Kamath
- Yun Liu (ORCID: https://orcid.org/0000-0003-4079-8275)
- Léonard Hussenot
- Bram Sterling
- Edouard Yvinec (ORCID: https://orcid.org/0000-0002-4318-612X)
- Shawn Xu
- Rory Pilgrim (ORCID: https://orcid.org/0009-0005-2582-8722)
- Johan Ferret
- Omar Sanseviero
- Olivier Bachem
- Ronnachai Jaroensri
- Vlad Feinberg
- Tatiana Matejovicova
- Tris Warkentin
- Victor Cotruta
- Daniel McDuff (ORCID: https://orcid.org/0000-0001-7313-0082)
- Jeremy Lai
- Richa Tiwari (ORCID: https://orcid.org/0000-0002-4003-1616)
- Avinatan Hassidim (ORCID: https://orcid.org/0000-0002-7034-7427)
- Chufan Gao
- Ramona Merhej
- Michelle Casbon
- Armand Joulin
- Shashir Reddy
- Morgane Rivière
- Gus Martins
- Alek Andreev
- Phoebe Kirk
- Fereshteh Mahvar (ORCID: https://orcid.org/0009-0009-6463-0451)
- Thomas Mesnard
- Ryan Brush
- Per Bjornsson
- Ines Mezerreg
- Madeleine Traverse
- Kejia Chen (ORCID: https://orcid.org/0009-0005-3779-3045)
- Dr. Anand Rao
- Susanna Maria Baby
- Shreya Pathak
- Kavi Goel
- Howard Yang
- Fayaz Jamil
- Howard Hu
- Catherine Kozlowski
- Sarah Perrin
- Lu Yang
- Preeti Singh
- Lin Yang
- Bashir Sadjad
Institutions
- Google (United States) (US)
Publication Details
- Journal
- Nature Medicine
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1038/s41591-026-04626-w
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00