Comparative evaluation of statistical learning methods for polygenic prediction in UK Biobank
Polygenic risk scores (PRS), which are calculated from genetic variants, provide a measure of an individual’s genetic predisposition to disease. They are widely used to predict disease risk, identify at-risk populations, and advance personalized medicine. However, PRS performance varies depending on genetic architecture, including heritability, the number of single nucleotide polymorphisms (SNPs), the proportion of causal SNPs, trait prevalence, and case–control imbalance. Despite their increasing use, methodological guidelines for selecting optimal PRS models tailored to different genetic architectures and trait characteristics remain limited. In this study, we evaluated and compared the performance of various PRS models, including infinitesimal, non-infinitesimal, and functional annotation-based models, such as LDpred2, PRS-CS, SBayesR, SBayesRC, PolyPred, and MegaPRS through comprehensive genome-wide simulations and real data analyses. Our simulations demonstrated that when the proportion of causal SNPs is low, non-infinitesimal models outperformed infinitesimal models, and LDpred-based methods generally achieved the highest accuracy. We further validated our findings using UK Biobank phenotypes, including height, body mass index, type 2 diabetes, and glaucoma, across both matched case–control analyses and the full UK Biobank cohort. In these real-data analyses, non-infinitesimal models also remained more accurate than infinitesimal models; PRS-CS and SBayesR performed comparably to, and for some traits better than, the LDpred-based methods among approaches without functional annotations, and functional annotation-based models further improved prediction accuracy, with SBayesRC and PolyPred generally achieving the highest predictive performance. Overall, our findings demonstrate that the optimal PRS method depends on the underlying genetic architecture of the target trait and model assumptions, and that functional genomic annotations may further improve prediction accuracy.
Authors
- Wonil Chung (ORCID: https://orcid.org/0000-0002-5766-6247)
- Seunghwan Park
Institutions
- Harvard University (US)
- Soongsil University (KR)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-06
- DOI
- https://doi.org/10.1038/s41598-026-70042-2
- Primary Topic
- Genetic Associations and Epidemiology
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Research Foundation of Korea