Markedly divergent performance of variant annotation methods for gene-level association testing

Abstract Machine learning-based annotation methods are increasingly used to assess the pathogenicity of genetic variants, but their performance at prioritizing variants for gene-level association testing remains poorly characterized. Here, to better understand and optimize for this use case, we assess variant annotations from five methods — CADD v1.6, CADD v1.7, AlphaMissense, ESM-1b, and GPN-MSA — as the basis of four primary gene-based tests and six annotation-level aggregation tests across 14 quantitative traits measured in up to 350,377 UK Biobank participants. Using a novel framework based on optimal transport, we quantify test calibration and power as relative measurements of how methods partition signal across labeled variant sets. These metrics reveal discordant performance characteristics across annotation methods: tests using CADD labels achieved the highest signal separation, while tests using AlphaMissense labels had the lowest calibration. Meanwhile, hits from tests using GPN-MSA labels were strongly enriched among genes with high evolutionary constraint (up to 5.8-fold, versus 2.1- to 2.8-fold using other methods), underscoring variability in how annotation methods assess pathogenicity across the genome. These differences in variant prioritization measurably shifted patterns of genetic discovery, even as tests had low genomic inflation and similar replication rates. Taken together, our results show that (1) no single combination of annotation method and statistical test is currently optimal for this task, (2) results from different methods and tests can and should be aggregated to increase study power, and (3) integrating the information content of variant and gene annotations remains a key open problem for rare variant association testing.

Authors

Publication Details

Journal
BMC Genomics
Published
2026-09-25
DOI
https://doi.org/10.1186/s12864-026-13379-2
Primary Topic
Genetic Associations and Epidemiology
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Markedly divergent performance of variant annotation methods for gene-level association testing

Flaviyan Jerome Irudayanathan, Hussein A. Hejase, Kipper Fletez‐Brant, Matthew W. Aguirre et al.
BMC Genomics
Genetic Associations and Epidemiology
article

Markedly divergent performance of variant annotation methods for gene-level association testing

Flaviyan Jerome Irudayanathan, Hussein A. Hejase, Kipper Fletez‐Brant, Matthew W. Aguirre, Mark I. McCarthy, Vipin K. Menon, Megan Crow, Sarah A. Pendergrass
article en

Abstract

Abstract Machine learning-based annotation methods are increasingly used to assess the pathogenicity of genetic variants, but their performance at prioritizing variants for gene-level association testing remains poorly characterized. Here, to better understand and optimize for this use case, we assess variant annotations from five methods — CADD v1.6, CADD v1.7, AlphaMissense, ESM-1b, and GPN-MSA — as the basis of four primary gene-based tests and six annotation-level aggregation tests across 14 quantitative traits measured in up to 350,377 UK Biobank participants. Using a novel framework based on optimal transport, we quantify test calibration and power as relative measurements of how methods partition signal across labeled variant sets. These metrics reveal discordant performance characteristics across annotation methods: tests using CADD labels achieved the highest signal separation, while tests using AlphaMissense labels had the lowest calibration. Meanwhile, hits from tests using GPN-MSA labels were strongly enriched among genes with high evolutionary constraint (up to 5.8-fold, versus 2.1- to 2.8-fold using other methods), underscoring variability in how annotation methods assess pathogenicity across the genome. These differences in variant prioritization measurably shifted patterns of genetic discovery, even as tests had low genomic inflation and similar replication rates. Taken together, our results show that (1) no single combination of annotation method and statistical test is currently optimal for this task, (2) results from different methods and tests can and should be aggregated to increase study power, and (3) integrating the information content of variant and gene annotations remains a key open problem for rare variant association testing.

BMC Genomics
Openalex Percentile: Top 12%
Genetic Associations and Epidemiology
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.