Comparative evaluation of statistical learning methods for polygenic prediction in UK Biobank

Polygenic risk scores (PRS), which are calculated from genetic variants, provide a measure of an individual’s genetic predisposition to disease. They are widely used to predict disease risk, identify at-risk populations, and advance personalized medicine. However, PRS performance varies depending on genetic architecture, including heritability, the number of single nucleotide polymorphisms (SNPs), the proportion of causal SNPs, trait prevalence, and case–control imbalance. Despite their increasing use, methodological guidelines for selecting optimal PRS models tailored to different genetic architectures and trait characteristics remain limited. In this study, we evaluated and compared the performance of various PRS models, including infinitesimal, non-infinitesimal, and functional annotation-based models, such as LDpred2, PRS-CS, SBayesR, SBayesRC, PolyPred, and MegaPRS through comprehensive genome-wide simulations and real data analyses. Our simulations demonstrated that when the proportion of causal SNPs is low, non-infinitesimal models outperformed infinitesimal models, and LDpred-based methods generally achieved the highest accuracy. We further validated our findings using UK Biobank phenotypes, including height, body mass index, type 2 diabetes, and glaucoma, across both matched case–control analyses and the full UK Biobank cohort. In these real-data analyses, non-infinitesimal models also remained more accurate than infinitesimal models; PRS-CS and SBayesR performed comparably to, and for some traits better than, the LDpred-based methods among approaches without functional annotations, and functional annotation-based models further improved prediction accuracy, with SBayesRC and PolyPred generally achieving the highest predictive performance. Overall, our findings demonstrate that the optimal PRS method depends on the underlying genetic architecture of the target trait and model assumptions, and that functional genomic annotations may further improve prediction accuracy.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-06
DOI
https://doi.org/10.1038/s41598-026-70042-2
Primary Topic
Genetic Associations and Epidemiology
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Comparative evaluation of statistical learning methods for polygenic prediction in UK Biobank

Wonil Chung, Seunghwan Park
Scientific Reports
Genetic Associations and Epidemiology
article

Comparative evaluation of statistical learning methods for polygenic prediction in UK Biobank

Wonil Chung, Seunghwan Park
article en

Abstract

Polygenic risk scores (PRS), which are calculated from genetic variants, provide a measure of an individual’s genetic predisposition to disease. They are widely used to predict disease risk, identify at-risk populations, and advance personalized medicine. However, PRS performance varies depending on genetic architecture, including heritability, the number of single nucleotide polymorphisms (SNPs), the proportion of causal SNPs, trait prevalence, and case–control imbalance. Despite their increasing use, methodological guidelines for selecting optimal PRS models tailored to different genetic architectures and trait characteristics remain limited. In this study, we evaluated and compared the performance of various PRS models, including infinitesimal, non-infinitesimal, and functional annotation-based models, such as LDpred2, PRS-CS, SBayesR, SBayesRC, PolyPred, and MegaPRS through comprehensive genome-wide simulations and real data analyses. Our simulations demonstrated that when the proportion of causal SNPs is low, non-infinitesimal models outperformed infinitesimal models, and LDpred-based methods generally achieved the highest accuracy. We further validated our findings using UK Biobank phenotypes, including height, body mass index, type 2 diabetes, and glaucoma, across both matched case–control analyses and the full UK Biobank cohort. In these real-data analyses, non-infinitesimal models also remained more accurate than infinitesimal models; PRS-CS and SBayesR performed comparably to, and for some traits better than, the LDpred-based methods among approaches without functional annotations, and functional annotation-based models further improved prediction accuracy, with SBayesRC and PolyPred generally achieving the highest predictive performance. Overall, our findings demonstrate that the optimal PRS method depends on the underlying genetic architecture of the target trait and model assumptions, and that functional genomic annotations may further improve prediction accuracy.

Scientific Reports
Harvard University (US), Soongsil University (KR)
National Research Foundation of Korea
Openalex Percentile: Top 11%
Genetic Associations and Epidemiology
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.