Surveillance-Driven Machine Learning for Prediction of Antimicrobial Susceptibility: An Explainable Modeling Framework using the Pfizer ATLAS Dataset (2004 - 2023)
Background Antimicrobial resistance (AMR) is a growing global public health threat, particularly in low- and middle-income countries (LMICs), where delayed antimicrobial susceptibility testing (AST) and limited diagnostic capacity complicate timely treatment and antimicrobial stewardship (AMS). Although large-scale surveillance programmes routinely collect antimicrobial susceptibility data, these datasets remain underutilised for predictive analytics. Methods We developed a surveillance-driven machine learning (ML) framework using isolate-level data from African sites participating in the Pfizer Antimicrobial Testing Leadership and Surveillance (ATLAS) programme between 2019 and 2023. Seven antibiotic-specific Extreme Gradient Boosting (XGBoost) models were developed to predict antimicrobial susceptibility using routinely collected microbiological, demographic, clinical, and geographical metadata. Model development included structured data preprocessing, RandomOverSampler-based class balancing restricted to the training data, hyperparameter optimisation using stratified cross-validation, and evaluation on a held-out test dataset. A rule-based MIC interpretation system and interactive dashboard were developed to demonstrate implementation of the analytical workflow. Results The antibiotic-specific models demonstrated moderate predictive performance, with test accuracies ranging from 57% to 76%. Ceftazidime-Avibactam achieved the highest test accuracy (76%), followed by Gentamicin (65%) and Imipenem (64%), while Amikacin showed the lowest performance (57%). Feature-importance analysis identified bacterial species as consistently among the most influential predictors. In contrast, bacterial family, country of isolate collection, specimen source, clinical specialty, and demographic characteristics contributed to varying degrees across antibiotics. Performance was generally stronger for the more frequently represented susceptible class than for intermediate and resistant isolates. Conclusion Routinely collected AMR surveillance data can support antibiotic-specific ML predictions without requiring genomic sequencing or detailed patient-level clinical information. This study provides a proof-of-concept framework for surveillance-driven predictive analytics that could complement conventional AMR surveillance and AMS. External validation, calibration, prospective clinical evaluation, and implementation studies are required before routine deployment.
Authors
- Frida Njeru
- David Gichohi
- Raphael Mutua (ORCID: https://orcid.org/0009-0003-8551-9134)
- Rahma O. Golicha
- Samuel K. Ndegwa (ORCID: https://orcid.org/0000-0002-3518-1142)
- Benson Kituku
Institutions
- Kenya Medical Research Institute (KE)
- Dedan Kimathi University of Technology (KE)
Publication Details
- Journal
- Wellcome Open Research
- Published
- 2026-09-21
- DOI
- https://doi.org/10.12688/wellcomeopenres.26477.2
- Primary Topic
- Bacterial Identification and Susceptibility Testing
- Type
- article
- Field-Weighted Citation Impact
- 0.00