PanDelos-plus: A parallel algorithm for computing sequence homology in pangenomic analysis
The identification of homologous gene families across multiple genomes is a central task in bacterial pangenomics traditionally requiring computationally demanding all-against-all comparisons. PanDelos addresses this challenge with an alignment-free and parameter-free approach based on k-mer profiles, combining high speed, ease of use, and competitive accuracy with state-of-the-art methods. However, the increasing availability of genomic data requires tools that can scale efficiently to larger datasets. To address this need, we present PanDelos-plus, a fully parallel, gene-centric redesign of PanDelos. The algorithm parallelizes the most computationally intensive phases (Best Hit detection and Bidirectional Best Hit extraction) through data decomposition and a thread pool strategy, while employing lightweight data structures to reduce memory usage. Benchmarks on synthetic datasets show that PanDelos-plus achieves up to 14x faster execution and reduces memory usage by up to 96%, while maintaining consistency with the original algorithm. These improvements allow the PanDelos methodology to be applied to population-scale comparative genomics, thus enabling more precise characterisation of pangenome structure and dynamics. PanDelos-plus is available at github.com/synbionics/PanDelos-plus.
Authors
- Vincenzo Bonnici (ORCID: https://orcid.org/0000-0002-1637-7545)
- Simone Colli
- Emiliano Maresi
Institutions
- University of Parma (IT)
Publication Details
- Journal
- PLoS Computational Biology
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1371/journal.pcbi.1014724
- Primary Topic
- Genomics and Phylogenetic Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Università degli Studi di Parma