ERMA: a harmonizing epicPCR data analysis tool
Abstract Background EpicPCR links functional genes to taxonomic markers at the single-cell level and has become an important method in microbial ecology and antimicrobial resistance research. However, the lack of missing standardized data analysis workflows limits reproducibility and comparability across studies. Existing analyses are often based on custom, unpublished scripts, complicating cross-study benchmarking and reuse. Results We present ERMA, an open-source, Snakemake-based pipeline that enables an automated, standardized, reproducible, and scalable epicPCR data analysis with minimal user input. ERMA supports both Illumina and ONT sequencing data and integrates quality control, database preparation, dual similarity searches against taxonomic and functional reference databases, results integration, filtering, and visualization. The pipeline is adaptable to a broad range of research questions beyond AMR due to modular database preparation and integration of specific UniRef queries. We evaluated ERMA using five publicly available epicPCR datasets covering diverse sample sources and sequencing depths. Across these datasets, ERMA achieved mean concordance rates of approximately 70–83% with the taxonomic results reported in the respective publications. In addition, validation using a defined five-member mock community spiked into clinical wastewater demonstrated full recovery of mock genera, with 85% of ERMA-detected genera also detected by independent 16S rDNA gene analysis of the same wastewater sample. Attrition analysis showed inconsistent filtering patterns across studies, reflecting dataset-specific quality differences. Conclusion While large-scale reference datasets for epicPCR remain limited, validation using both publicly available studies and a defined mock community shows strong, interpretable performance across heterogeneous datasets. By providing a transparent, modular, and reproducible workflow, ERMA contributes toward harmonized epicPCR data analysis and establishes a foundation for future methodological standardization and benchmarking efforts as reference datasets improve.
Authors
- Jan Kehrmann (ORCID: https://orcid.org/0000-0002-1494-9191)
- Folker Meyer (ORCID: https://orcid.org/0000-0003-1112-2284)
- Ivana Kraiselburd (ORCID: https://orcid.org/0000-0002-9677-6094)
- Adrian Dörr (ORCID: https://orcid.org/0009-0000-2526-0384)
- Svjetlana Dekić Rozman
- Marko Virta
- Jan Buer
- Lina Viktoria Duncker
Publication Details
- Journal
- BMC Genomics
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1186/s12864-026-13418-y
- Primary Topic
- Genomics and Phylogenetic Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00