ERMA: a harmonizing epicPCR data analysis tool

Abstract Background EpicPCR links functional genes to taxonomic markers at the single-cell level and has become an important method in microbial ecology and antimicrobial resistance research. However, the lack of missing standardized data analysis workflows limits reproducibility and comparability across studies. Existing analyses are often based on custom, unpublished scripts, complicating cross-study benchmarking and reuse. Results We present ERMA, an open-source, Snakemake-based pipeline that enables an automated, standardized, reproducible, and scalable epicPCR data analysis with minimal user input. ERMA supports both Illumina and ONT sequencing data and integrates quality control, database preparation, dual similarity searches against taxonomic and functional reference databases, results integration, filtering, and visualization. The pipeline is adaptable to a broad range of research questions beyond AMR due to modular database preparation and integration of specific UniRef queries. We evaluated ERMA using five publicly available epicPCR datasets covering diverse sample sources and sequencing depths. Across these datasets, ERMA achieved mean concordance rates of approximately 70–83% with the taxonomic results reported in the respective publications. In addition, validation using a defined five-member mock community spiked into clinical wastewater demonstrated full recovery of mock genera, with 85% of ERMA-detected genera also detected by independent 16S rDNA gene analysis of the same wastewater sample. Attrition analysis showed inconsistent filtering patterns across studies, reflecting dataset-specific quality differences. Conclusion While large-scale reference datasets for epicPCR remain limited, validation using both publicly available studies and a defined mock community shows strong, interpretable performance across heterogeneous datasets. By providing a transparent, modular, and reproducible workflow, ERMA contributes toward harmonized epicPCR data analysis and establishes a foundation for future methodological standardization and benchmarking efforts as reference datasets improve.

Authors

Publication Details

Journal
BMC Genomics
Published
2026-10-09
DOI
https://doi.org/10.1186/s12864-026-13418-y
Primary Topic
Genomics and Phylogenetic Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

ERMA: a harmonizing epicPCR data analysis tool

Jan Kehrmann, Folker Meyer, Ivana Kraiselburd, Adrian Dörr et al.
BMC Genomics
Genomics and Phylogenetic Studies
article

ERMA: a harmonizing epicPCR data analysis tool

Jan Kehrmann, Folker Meyer, Ivana Kraiselburd, Adrian Dörr, Svjetlana Dekić Rozman, Marko Virta, Jan Buer, Lina Viktoria Duncker
article en

Abstract

Abstract Background EpicPCR links functional genes to taxonomic markers at the single-cell level and has become an important method in microbial ecology and antimicrobial resistance research. However, the lack of missing standardized data analysis workflows limits reproducibility and comparability across studies. Existing analyses are often based on custom, unpublished scripts, complicating cross-study benchmarking and reuse. Results We present ERMA, an open-source, Snakemake-based pipeline that enables an automated, standardized, reproducible, and scalable epicPCR data analysis with minimal user input. ERMA supports both Illumina and ONT sequencing data and integrates quality control, database preparation, dual similarity searches against taxonomic and functional reference databases, results integration, filtering, and visualization. The pipeline is adaptable to a broad range of research questions beyond AMR due to modular database preparation and integration of specific UniRef queries. We evaluated ERMA using five publicly available epicPCR datasets covering diverse sample sources and sequencing depths. Across these datasets, ERMA achieved mean concordance rates of approximately 70–83% with the taxonomic results reported in the respective publications. In addition, validation using a defined five-member mock community spiked into clinical wastewater demonstrated full recovery of mock genera, with 85% of ERMA-detected genera also detected by independent 16S rDNA gene analysis of the same wastewater sample. Attrition analysis showed inconsistent filtering patterns across studies, reflecting dataset-specific quality differences. Conclusion While large-scale reference datasets for epicPCR remain limited, validation using both publicly available studies and a defined mock community shows strong, interpretable performance across heterogeneous datasets. By providing a transparent, modular, and reproducible workflow, ERMA contributes toward harmonized epicPCR data analysis and establishes a foundation for future methodological standardization and benchmarking efforts as reference datasets improve.

BMC Genomics
Openalex Percentile: Top 23%
Genomics and Phylogenetic Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.