Apparent Performance of Antimicrobial Resistance Phenotype Prediction Depends on the Definition of an Unseen Observation

Most machine-learning studies of antimicrobial resistance (AMR) evaluate their models by splitting rows at random into training and test sets. That keeps individual rows apart, but it does nothing to stop the same species, antibiotic, species–antibiotic pairing or genome from turning up on both sides of the split. We tested how much this choice actually matters, using a genome-level AMR dataset from BV-BRC (801 labelled ob servations spanning 5 species, 41 genomes, 46 antibiotics and 82 species–antibiotic combinations). Four classifiers — logistic regression, random forest, gradient boosting and XGBoost — were run under fourevaluation regimes, each defining "unseen" differently: a random row-level split; a group k-fold holding out species–antibiotic combinations; a group k-fold holding out genomes; and a leave-one-species-out design. Alongside these four regimes we built relationship-based majority-vote baselines, audited each split for training–test overlap, ran feature ablations, tested robustness across model seeds, permuted labels and checked for label conflicts within relationships. The pattern that emerged was consistent across classifiers. Gradient boosting reached 0.824 accuracy under random row-level splitting, dropped to 0.628 once species–antibiotic combinations were held out, climbed back to 0.824 when only genomes were withheld, and fell to 0.406 (pooled across all five held-out species; 0.438 by unweighted fold-mean) under leave-one-species-out. A simple majority-label lookup based on Species+Antibiotic already reached 0.860 in-sample accuracy, and a Genome+Antibiotic lookup reached 0.985 — and species–antibiotic overlap stayed at 100% even in the unseen-genome regime, which helps explain why that regime performed almost as well as the random split. Permutation testing showed observed accuracy beating a shuffled-label null under the first three regimes (all P = 0.005), but not under leave-one-out, where observed accuracy failed to exceed the null distribution at all (pooled statistic, P = 1.000). Taken together, these results show that the apparent performance of AMR phenotype prediction depends heavily on what a study defines as an unseen observation. Random row-level splitting produces an optimistic estimate compared with regimes that actually control species–antibiotic or species-level overlap, and much of the accuracy achieved under the more permissive regimes looks less like genuine generalisation and more like a model exploiting relationships it has already seen. AMR prediction studies should state explicitly which biological entities and relationships are prevented from overlapping between training and test data.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-26
DOI
https://doi.org/10.5281/zenodo.22976451
Primary Topic
Antibiotic Use and Resistance
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Apparent Performance of Antimicrobial Resistance Phenotype Prediction Depends on the Definition of an Unseen Observation

Mohammad Ayesha Summaiyya, Kavya Sudha Vadlamudi
Zenodo (CERN European Organization for Nuclear Research)
Antibiotic Use and Resistance
preprint

Apparent Performance of Antimicrobial Resistance Phenotype Prediction Depends on the Definition of an Unseen Observation

Mohammad Ayesha Summaiyya, Kavya Sudha Vadlamudi
preprint en

Abstract

Most machine-learning studies of antimicrobial resistance (AMR) evaluate their models by splitting rows at random into training and test sets. That keeps individual rows apart, but it does nothing to stop the same species, antibiotic, species–antibiotic pairing or genome from turning up on both sides of the split. We tested how much this choice actually matters, using a genome-level AMR dataset from BV-BRC (801 labelled ob servations spanning 5 species, 41 genomes, 46 antibiotics and 82 species–antibiotic combinations). Four classifiers — logistic regression, random forest, gradient boosting and XGBoost — were run under fourevaluation regimes, each defining "unseen" differently: a random row-level split; a group k-fold holding out species–antibiotic combinations; a group k-fold holding out genomes; and a leave-one-species-out design. Alongside these four regimes we built relationship-based majority-vote baselines, audited each split for training–test overlap, ran feature ablations, tested robustness across model seeds, permuted labels and checked for label conflicts within relationships. The pattern that emerged was consistent across classifiers. Gradient boosting reached 0.824 accuracy under random row-level splitting, dropped to 0.628 once species–antibiotic combinations were held out, climbed back to 0.824 when only genomes were withheld, and fell to 0.406 (pooled across all five held-out species; 0.438 by unweighted fold-mean) under leave-one-species-out. A simple majority-label lookup based on Species+Antibiotic already reached 0.860 in-sample accuracy, and a Genome+Antibiotic lookup reached 0.985 — and species–antibiotic overlap stayed at 100% even in the unseen-genome regime, which helps explain why that regime performed almost as well as the random split. Permutation testing showed observed accuracy beating a shuffled-label null under the first three regimes (all P = 0.005), but not under leave-one-out, where observed accuracy failed to exceed the null distribution at all (pooled statistic, P = 1.000). Taken together, these results show that the apparent performance of AMR phenotype prediction depends heavily on what a study defines as an unseen observation. Random row-level splitting produces an optimistic estimate compared with regimes that actually control species–antibiotic or species-level overlap, and much of the accuracy achieved under the more permissive regimes looks less like genuine generalisation and more like a model exploiting relationships it has already seen. AMR prediction studies should state explicitly which biological entities and relationships are prevented from overlapping between training and test data.

Zenodo (CERN European Organization for Nuclear Research)
Life in Land
Antibiotic Use and Resistance
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.