Generalizable CT vision-language modeling for population health and disease risk
Abstract Vision-language foundation models (VLMs) for computed tomography (CT) are emerging tools that learn generalizable representations from large-scale clinical imaging data. While these models can predict task-specific labels, the extent to which their representations capture the clinical, physiological, and longitudinal variation of real-world patient populations remains unclear. We introduce Percival, a CT-native VLM trained on more than 400,000 CT-report pairs from the Penn Medicine BioBank using a dual-encoder symmetric contrastive framework. Across over 20,000 held-out participants, Percival’s latent space aligns with demographic, physiological, and laboratory variation, and supports phenome-wide associations across the electronic health record. We evaluate Percival against alternative foundation-model paradigms, including vision-only contrastive and multi-organ segmentation, demonstrating that vision-language pretraining captures clinical information not fully accessible to vision-only alternatives across disease classification and longitudinal risk modeling. Together, these findings indicate that CT-VLMs uncover latent structure aligned with clinical, physiological, and prognostic variation across the disease-prevalence spectrum.
Authors
- Julio A. Chirinos (ORCID: https://orcid.org/0000-0001-9035-5670)
- Bojian Hou (ORCID: https://orcid.org/0000-0002-3894-4547)
- Colleen Morse
- Anurag Verma (ORCID: https://orcid.org/0000-0002-5063-9107)
- James Gee
- Eric R. Eaton (ORCID: https://orcid.org/0000-0002-5689-2234)
- Hersh Sagreiya (ORCID: https://orcid.org/0000-0002-2909-6793)
- Farouk Dako (ORCID: https://orcid.org/0000-0003-4765-9358)
- Cameron Beeche (ORCID: https://orcid.org/0000-0002-5781-8810)
- Christos A. Davatzikos (ORCID: https://orcid.org/0000-0002-1025-8561)
- Joonghyun Kim (ORCID: https://orcid.org/0000-0002-5618-5249)
- Walter R T Witschey (ORCID: https://orcid.org/0000-0003-1669-2120)
- Hamed Tavolinejad (ORCID: https://orcid.org/0000-0002-0244-5914)
- Daniel James Rader (ORCID: https://orcid.org/0000-0002-9245-9876)
- Rohan Shad (ORCID: https://orcid.org/0000-0002-0453-9041)
- Marylyn DeRiggi Ritchie (ORCID: https://orcid.org/0000-0002-1208-1720)
- Scott M. Damrauer (ORCID: https://orcid.org/0000-0001-8009-1632)
- Charles E. Kahn (ORCID: https://orcid.org/0000-0002-6654-7434)
- Li Shen (ORCID: https://orcid.org/0000-0002-5443-0503)
- Jeffrey T. Duda
- Penn Medicine Biobank
- Jessie Dong (ORCID: https://orcid.org/0009-0001-5307-6610)
- Bingxin Zhao
- Rakesh Sharma
- Gengwei Zhang
- Qi Long
- Tianlong Chen
Institutions
- University of North Carolina at Chapel Hill (US)
- Hospital of the University of Pennsylvania (US)
- Philadelphia VA Medical Center (US)
- University of Pennsylvania (US)
Publication Details
- Journal
- npj Digital Medicine
- Published
- 2026-09-29
- DOI
- https://doi.org/10.1038/s41746-026-03257-2
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00