Known QAC determinants outperform genome language model embeddings in a leakage-aware public Listeria benzalkonium chloride tolerance benchmark
Abstract Whole-genome sequencing and pretrained DNA models may help prioritize disinfectant-tolerance hypotheses, but phenotype prediction requires measured, isolate-linked labels and evaluation that keeps related lineages apart. We compared conventional sequence features, numerical DNA representations generated by DNABERT2, and known quaternary ammonium compound (QAC) determinants in a public benchmark of 197 Listeria monocytogenes isolates with benzalkonium chloride MIC data. Models were evaluated across five repeated lineage-grouped train, development, and test splits. The known-QAC determinant rule showed the highest balanced accuracy (mean ± SD, 0.961 ± 0.057) and ROC AUC (0.961 ± 0.057). A model combining k-mer, DNABERT2, and QAC scores reached balanced accuracy 0.915 ± 0.023, whereas the best DNABERT2-based classifier reached 0.570 ± 0.110; sparse k-mer models did not exceed the dummy baseline. Public non-Listeria studies lacked the row-level phenotype-genome linkage required for supervised modelling. An exploratory bioinformatic screen therefore identified only candidate regions for future testing, not validated tolerance predictions. Because the benchmark contained only 197 publicly linkable isolates, these findings are hypothesis-generating and require confirmation in larger, independent phenotype-linked cohorts. Within these limits, interpretable determinants were the primary validated signal and DNA-model scores were secondary evidence.
Authors
- Carlos Victor Montefusco-Pereira (ORCID: https://orcid.org/0000-0003-4167-4653)
Institutions
- Artificial Intelligence in Medicine (Canada) (CA)
Publication Details
- Journal
- FEMS Microbiology Letters
- Published
- 2026-09-26
- DOI
- https://doi.org/10.1093/femsle/fnag112
- Primary Topic
- Environmental Chemistry and Analysis
- Type
- article
- Field-Weighted Citation Impact
- 0.00