Generalizable CT vision-language modeling for population health and disease risk

Abstract Vision-language foundation models (VLMs) for computed tomography (CT) are emerging tools that learn generalizable representations from large-scale clinical imaging data. While these models can predict task-specific labels, the extent to which their representations capture the clinical, physiological, and longitudinal variation of real-world patient populations remains unclear. We introduce Percival, a CT-native VLM trained on more than 400,000 CT-report pairs from the Penn Medicine BioBank using a dual-encoder symmetric contrastive framework. Across over 20,000 held-out participants, Percival’s latent space aligns with demographic, physiological, and laboratory variation, and supports phenome-wide associations across the electronic health record. We evaluate Percival against alternative foundation-model paradigms, including vision-only contrastive and multi-organ segmentation, demonstrating that vision-language pretraining captures clinical information not fully accessible to vision-only alternatives across disease classification and longitudinal risk modeling. Together, these findings indicate that CT-VLMs uncover latent structure aligned with clinical, physiological, and prognostic variation across the disease-prevalence spectrum.

Authors

Institutions

Publication Details

Journal
npj Digital Medicine
Published
2026-09-29
DOI
https://doi.org/10.1038/s41746-026-03257-2
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Generalizable CT vision-language modeling for population health and disease risk

Julio A. Chirinos, Bojian Hou, Colleen Morse, Anurag Verma et al.
npj Digital Medicine
Artificial Intelligence in Healthcare and Education
article

Generalizable CT vision-language modeling for population health and disease risk

Julio A. Chirinos, Bojian Hou, Colleen Morse, Anurag Verma, James Gee, Eric R. Eaton, Hersh Sagreiya, Farouk Dako, Cameron Beeche, Christos A. Davatzikos, Joonghyun Kim, Walter R T Witschey, Hamed Tavolinejad, Daniel James Rader, Rohan Shad, Marylyn DeRiggi Ritchie, Scott M. Damrauer, Charles E. Kahn, Li Shen, Jeffrey T. Duda, Penn Medicine Biobank, Jessie Dong, Bingxin Zhao, Rakesh Sharma, Gengwei Zhang, Qi Long, Tianlong Chen
article en

Abstract

Abstract Vision-language foundation models (VLMs) for computed tomography (CT) are emerging tools that learn generalizable representations from large-scale clinical imaging data. While these models can predict task-specific labels, the extent to which their representations capture the clinical, physiological, and longitudinal variation of real-world patient populations remains unclear. We introduce Percival, a CT-native VLM trained on more than 400,000 CT-report pairs from the Penn Medicine BioBank using a dual-encoder symmetric contrastive framework. Across over 20,000 held-out participants, Percival’s latent space aligns with demographic, physiological, and laboratory variation, and supports phenome-wide associations across the electronic health record. We evaluate Percival against alternative foundation-model paradigms, including vision-only contrastive and multi-organ segmentation, demonstrating that vision-language pretraining captures clinical information not fully accessible to vision-only alternatives across disease classification and longitudinal risk modeling. Together, these findings indicate that CT-VLMs uncover latent structure aligned with clinical, physiological, and prognostic variation across the disease-prevalence spectrum.

npj Digital Medicine
University of North Carolina at Chapel Hill (US), Hospital of the University of Pennsylvania (US), Philadelphia VA Medical Center (US), University of Pennsylvania (US)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.