DistPCA: Tera-Scale Genomic PCA via Out-of-Core Distributed Parallelism

Abstract Motivation Principal Component Analysis (PCA) is a core component of human genomic pipelines, widely used for population structure inference, ancestry analysis, and quality control in genome-wide association studies. Over the past decade, the increasing scale of genomic datasets has pushed PCA beyond the limits of in-core computation, motivating the adoption of out-of-core methods. However, existing approaches remain limited to single-node infrastructure and fail to exploit available parallelism in data fetching and preprocessing. As cohorts grow toward next-generation biobank scale, these limitations create a scalability barrier that makes routine PCA increasingly impractical for large-scale genetic and population analysis. Results We introduce DistPCA, a first distributed out-of-core framework for tera-scale genomic PCA, implemented as a high-performance C ++ software package that scales from single-node to multi-node computing environments. Built on top of Message Passing Interface (MPI), DistPCA employs hybrid multi-level data parallelism across the entire PCA pipeline, including data fetching, preprocessing, and numerical computation. Extensive evaluation on real and synthetic datasets demonstrates near-linear scalability, reducing wall-clock time by more than 80% compared with current state-of-the-art methods (from 12.1h to 2.3h). These results establish DistPCA as a robust solution for scalable routine population structure analysis at next-generation biobank scale. Availability and Implementation The source code and documentation of DistPCA are available on GitHub at https://github.com/CEID-HPCLAB/DistPCA and on Zenodo at https://doi.org/10.5281/zenodo.20392865.

Authors

Institutions

Publication Details

Journal
Bioinformatics Advances
Published
2026-10-06
DOI
https://doi.org/10.1093/bioadv/vbag303
Primary Topic
Parallel Computing and Optimization Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

DistPCA: Tera-Scale Genomic PCA via Out-of-Core Distributed Parallelism

Eugenia-Maria Kontopoulou, Efstratios Gallopoulos, Panagiotis E. Hadjidoukas, Argiris Sofotasios et al.
Bioinformatics Advances
Parallel Computing and Optimization Techniques
article

DistPCA: Tera-Scale Genomic PCA via Out-of-Core Distributed Parallelism

Eugenia-Maria Kontopoulou, Efstratios Gallopoulos, Panagiotis E. Hadjidoukas, Argiris Sofotasios, Georgios Mermigkis
article en

Abstract

Abstract Motivation Principal Component Analysis (PCA) is a core component of human genomic pipelines, widely used for population structure inference, ancestry analysis, and quality control in genome-wide association studies. Over the past decade, the increasing scale of genomic datasets has pushed PCA beyond the limits of in-core computation, motivating the adoption of out-of-core methods. However, existing approaches remain limited to single-node infrastructure and fail to exploit available parallelism in data fetching and preprocessing. As cohorts grow toward next-generation biobank scale, these limitations create a scalability barrier that makes routine PCA increasingly impractical for large-scale genetic and population analysis. Results We introduce DistPCA, a first distributed out-of-core framework for tera-scale genomic PCA, implemented as a high-performance C ++ software package that scales from single-node to multi-node computing environments. Built on top of Message Passing Interface (MPI), DistPCA employs hybrid multi-level data parallelism across the entire PCA pipeline, including data fetching, preprocessing, and numerical computation. Extensive evaluation on real and synthetic datasets demonstrates near-linear scalability, reducing wall-clock time by more than 80% compared with current state-of-the-art methods (from 12.1h to 2.3h). These results establish DistPCA as a robust solution for scalable routine population structure analysis at next-generation biobank scale. Availability and Implementation The source code and documentation of DistPCA are available on GitHub at https://github.com/CEID-HPCLAB/DistPCA and on Zenodo at https://doi.org/10.5281/zenodo.20392865.

Bioinformatics Advances
Computer Technology Institute and Press “DIOPHANTUS” (GR), University of Patras (GR)
Openalex Percentile: Top 8%
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.