Apparent Performance of Antimicrobial Resistance Phenotype Prediction Depends on the Definition of an Unseen Observation
Most machine-learning studies of antimicrobial resistance (AMR) evaluate their models by splitting rows at random into training and test sets. That keeps individual rows apart, but it does nothing to stop the same species, antibiotic, species–antibiotic pairing or genome from turning up on both sides of the split. We tested how much this choice actually matters, using a genome-level AMR dataset from BV-BRC (801 labelled ob servations spanning 5 species, 41 genomes, 46 antibiotics and 82 species–antibiotic combinations). Four classifiers — logistic regression, random forest, gradient boosting and XGBoost — were run under fourevaluation regimes, each defining "unseen" differently: a random row-level split; a group k-fold holding out species–antibiotic combinations; a group k-fold holding out genomes; and a leave-one-species-out design. Alongside these four regimes we built relationship-based majority-vote baselines, audited each split for training–test overlap, ran feature ablations, tested robustness across model seeds, permuted labels and checked for label conflicts within relationships. The pattern that emerged was consistent across classifiers. Gradient boosting reached 0.824 accuracy under random row-level splitting, dropped to 0.628 once species–antibiotic combinations were held out, climbed back to 0.824 when only genomes were withheld, and fell to 0.406 (pooled across all five held-out species; 0.438 by unweighted fold-mean) under leave-one-species-out. A simple majority-label lookup based on Species+Antibiotic already reached 0.860 in-sample accuracy, and a Genome+Antibiotic lookup reached 0.985 — and species–antibiotic overlap stayed at 100% even in the unseen-genome regime, which helps explain why that regime performed almost as well as the random split. Permutation testing showed observed accuracy beating a shuffled-label null under the first three regimes (all P = 0.005), but not under leave-one-out, where observed accuracy failed to exceed the null distribution at all (pooled statistic, P = 1.000). Taken together, these results show that the apparent performance of AMR phenotype prediction depends heavily on what a study defines as an unseen observation. Random row-level splitting produces an optimistic estimate compared with regimes that actually control species–antibiotic or species-level overlap, and much of the accuracy achieved under the more permissive regimes looks less like genuine generalisation and more like a model exploiting relationships it has already seen. AMR prediction studies should state explicitly which biological entities and relationships are prevented from overlapping between training and test data.
Authors
- Mohammad Ayesha Summaiyya
- Kavya Sudha Vadlamudi (ORCID: https://orcid.org/0009-0009-8556-9979)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-26
- DOI
- https://doi.org/10.5281/zenodo.22976451
- Primary Topic
- Antibiotic Use and Resistance
- Type
- preprint