Wild-AC: A fast, index-free multi-pattern string matching algorithm with wildcard support for proteomics
Abstract Background Efficient multi-pattern string matching with wildcard support is essential in proteomics, where peptides must be mapped to protein databases containing ambiguous amino acids (e.g., B , Z , X ). Existing approaches either require preprocessing (e.g., FM-index) or fail to handle wildcards in the text. Results We present Wild-AC, a modified Aho-Corasick algorithm that supports wildcards in the text by branching the search into parallel ‘scout’ paths, one per possible wildcard representation, while the unmodified primary search continues unaffected in the common, wildcard-free case. Wild-AC outperforms the FM-index (wildcard case) and matches or exceeds Wu-Manber (exact search) in speed for realistic proteomics workloads, i.e., 1000–500 000 peptide patterns (average length ∼18 amino acids) searched against protein databases (texts) of 3 × 10 6 –2.1 × 10 8 characters. On databases masked to a wildcard rate of 5%, Wild-AC retains its advantage for large peptide sets, while the FM-index becomes preferable for small ones. It requires no index, scales well with pattern count, and supports multi-threading. The C++ implementation is open source at https://github.com/Wild-AC/Wild-AC . Conclusion Wild-AC extends the Aho-Corasick algorithm to support wildcards and outperforms FM-Index and Wu-Manber in the task of peptide-protein mapping while considering ambiguous amino acids.
Authors
- Chris Bielow (ORCID: https://orcid.org/0000-0001-5756-3988)
Institutions
- Freie Universität Berlin (DE)
Publication Details
- Journal
- BMC Bioinformatics
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1186/s12859-026-06686-8
- Primary Topic
- Algorithms and Data Compression
- Type
- article
- Field-Weighted Citation Impact
- 0.00