pileup-hi: an ultra-high-throughput, customizable alignment pileup program for large datasets

Abstract Motivation Recent advancements in next-generation sequencing have combined massive data throughput with high read accuracy to facilitate large-scale, rapid, and sensitive analysis of genetic variation. High-depth, high-quality platforms such as the Illumina NovaSeq X and Ultima Genomics UG 100 are increasingly used in time-sensitive clinical contexts such as cancer screening to detect mutations as low as 0.01% without sequence error correction. This process usually involves the bioinformatic construction of a pileup, or a list of nucleotides aligned to one or more positions in a sequence alignment, to identify variants. Efficient pileup software is required to process large-scale sequencing datasets rapidly to allow for timely clinical decision-making. While foundational to many analytical pipelines, the de facto standard pileup software samtools mpileup faces scalability challenges with larger datasets and is restricted to one output format. Results We present pileup-hi, a multi-threaded pileup engine that is scalable to alignments containing billions of reads and extendable to support custom output formats. When set to emit the default mpileup format, pileup-hi is up to 13x faster than samtools mpileup, 3.2x faster than sambamba mpileup, and up to 8x faster than perbase base-depth. The default output of pileup-hi has binary equivalence to the output of samtools mpileup across benchmark files and a variety of samtools regression tests. We present a new pileup-derived format that provides depth-invariant data storage proportional only to the reference genome length and number of unique indels. Availability and implementation Pileup-hi is implemented in the Rust programming language and distributed as open-source code and precompiled binaries. Source code and installation instructions can be found at https://github.com/greninger-lab/pileup-hi. Supplementary information Supplementary data are available at Bioinformatics online.

Authors

Institutions

Publication Details

Journal
Bioinformatics
Published
2026-09-28
DOI
https://doi.org/10.1093/bioinformatics/btag721
Primary Topic
vaccines and immunoinformatics approaches
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

pileup-hi: an ultra-high-throughput, customizable alignment pileup program for large datasets

A L Greninger, E Piliper
Bioinformatics
vaccines and immunoinformatics approaches
article

pileup-hi: an ultra-high-throughput, customizable alignment pileup program for large datasets

A L Greninger, E Piliper
article en

Abstract

Abstract Motivation Recent advancements in next-generation sequencing have combined massive data throughput with high read accuracy to facilitate large-scale, rapid, and sensitive analysis of genetic variation. High-depth, high-quality platforms such as the Illumina NovaSeq X and Ultima Genomics UG 100 are increasingly used in time-sensitive clinical contexts such as cancer screening to detect mutations as low as 0.01% without sequence error correction. This process usually involves the bioinformatic construction of a pileup, or a list of nucleotides aligned to one or more positions in a sequence alignment, to identify variants. Efficient pileup software is required to process large-scale sequencing datasets rapidly to allow for timely clinical decision-making. While foundational to many analytical pipelines, the de facto standard pileup software samtools mpileup faces scalability challenges with larger datasets and is restricted to one output format. Results We present pileup-hi, a multi-threaded pileup engine that is scalable to alignments containing billions of reads and extendable to support custom output formats. When set to emit the default mpileup format, pileup-hi is up to 13x faster than samtools mpileup, 3.2x faster than sambamba mpileup, and up to 8x faster than perbase base-depth. The default output of pileup-hi has binary equivalence to the output of samtools mpileup across benchmark files and a variety of samtools regression tests. We present a new pileup-derived format that provides depth-invariant data storage proportional only to the reference genome length and number of unique indels. Availability and implementation Pileup-hi is implemented in the Rust programming language and distributed as open-source code and precompiled binaries. Source code and installation instructions can be found at https://github.com/greninger-lab/pileup-hi. Supplementary information Supplementary data are available at Bioinformatics online.

Bioinformatics
Cape Town HVTN Immunology Laboratory / Hutchinson Centre Research Institute of South Africa (ZA), Infectious Disease Research Institute (US), University of Washington Medical Center (US), Fred Hutch Cancer Center (US)
Peace, Justice and strong institutions
Openalex Percentile: Top 20%
vaccines and immunoinformatics approaches
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.