Surveillance-Driven Machine Learning for Prediction of Antimicrobial Susceptibility: An Explainable Modeling Framework using the Pfizer ATLAS Dataset (2004 - 2023)

Background Antimicrobial resistance (AMR) is a growing global public health threat, particularly in low- and middle-income countries (LMICs), where delayed antimicrobial susceptibility testing (AST) and limited diagnostic capacity complicate timely treatment and antimicrobial stewardship (AMS). Although large-scale surveillance programmes routinely collect antimicrobial susceptibility data, these datasets remain underutilised for predictive analytics. Methods We developed a surveillance-driven machine learning (ML) framework using isolate-level data from African sites participating in the Pfizer Antimicrobial Testing Leadership and Surveillance (ATLAS) programme between 2019 and 2023. Seven antibiotic-specific Extreme Gradient Boosting (XGBoost) models were developed to predict antimicrobial susceptibility using routinely collected microbiological, demographic, clinical, and geographical metadata. Model development included structured data preprocessing, RandomOverSampler-based class balancing restricted to the training data, hyperparameter optimisation using stratified cross-validation, and evaluation on a held-out test dataset. A rule-based MIC interpretation system and interactive dashboard were developed to demonstrate implementation of the analytical workflow. Results The antibiotic-specific models demonstrated moderate predictive performance, with test accuracies ranging from 57% to 76%. Ceftazidime-Avibactam achieved the highest test accuracy (76%), followed by Gentamicin (65%) and Imipenem (64%), while Amikacin showed the lowest performance (57%). Feature-importance analysis identified bacterial species as consistently among the most influential predictors. In contrast, bacterial family, country of isolate collection, specimen source, clinical specialty, and demographic characteristics contributed to varying degrees across antibiotics. Performance was generally stronger for the more frequently represented susceptible class than for intermediate and resistant isolates. Conclusion Routinely collected AMR surveillance data can support antibiotic-specific ML predictions without requiring genomic sequencing or detailed patient-level clinical information. This study provides a proof-of-concept framework for surveillance-driven predictive analytics that could complement conventional AMR surveillance and AMS. External validation, calibration, prospective clinical evaluation, and implementation studies are required before routine deployment.

Authors

Institutions

Publication Details

Journal
Wellcome Open Research
Published
2026-09-21
DOI
https://doi.org/10.12688/wellcomeopenres.26477.2
Primary Topic
Bacterial Identification and Susceptibility Testing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Surveillance-Driven Machine Learning for Prediction of Antimicrobial Susceptibility: An Explainable Modeling Framework using the Pfizer ATLAS Dataset (2004 - 2023)

Frida Njeru, David Gichohi, Raphael Mutua, Rahma O. Golicha et al.
Wellcome Open Research
Bacterial Identification and Susceptibility Testing
article

Surveillance-Driven Machine Learning for Prediction of Antimicrobial Susceptibility: An Explainable Modeling Framework using the Pfizer ATLAS Dataset (2004 - 2023)

Frida Njeru, David Gichohi, Raphael Mutua, Rahma O. Golicha, Samuel K. Ndegwa, Benson Kituku
article en

Abstract

Background Antimicrobial resistance (AMR) is a growing global public health threat, particularly in low- and middle-income countries (LMICs), where delayed antimicrobial susceptibility testing (AST) and limited diagnostic capacity complicate timely treatment and antimicrobial stewardship (AMS). Although large-scale surveillance programmes routinely collect antimicrobial susceptibility data, these datasets remain underutilised for predictive analytics. Methods We developed a surveillance-driven machine learning (ML) framework using isolate-level data from African sites participating in the Pfizer Antimicrobial Testing Leadership and Surveillance (ATLAS) programme between 2019 and 2023. Seven antibiotic-specific Extreme Gradient Boosting (XGBoost) models were developed to predict antimicrobial susceptibility using routinely collected microbiological, demographic, clinical, and geographical metadata. Model development included structured data preprocessing, RandomOverSampler-based class balancing restricted to the training data, hyperparameter optimisation using stratified cross-validation, and evaluation on a held-out test dataset. A rule-based MIC interpretation system and interactive dashboard were developed to demonstrate implementation of the analytical workflow. Results The antibiotic-specific models demonstrated moderate predictive performance, with test accuracies ranging from 57% to 76%. Ceftazidime-Avibactam achieved the highest test accuracy (76%), followed by Gentamicin (65%) and Imipenem (64%), while Amikacin showed the lowest performance (57%). Feature-importance analysis identified bacterial species as consistently among the most influential predictors. In contrast, bacterial family, country of isolate collection, specimen source, clinical specialty, and demographic characteristics contributed to varying degrees across antibiotics. Performance was generally stronger for the more frequently represented susceptible class than for intermediate and resistant isolates. Conclusion Routinely collected AMR surveillance data can support antibiotic-specific ML predictions without requiring genomic sequencing or detailed patient-level clinical information. This study provides a proof-of-concept framework for surveillance-driven predictive analytics that could complement conventional AMR surveillance and AMS. External validation, calibration, prospective clinical evaluation, and implementation studies are required before routine deployment.

Wellcome Open ResearchVol. 11
Kenya Medical Research Institute (KE), Dedan Kimathi University of Technology (KE)
No poverty
Openalex Percentile: Top 14%
Bacterial Identification and Susceptibility Testing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.