Comparison of Machine Learning Classifiers for Predicting Reported Antibiotic Treatment After Recent Fever/Cough Among Under‐Five Children: A Pooled Analysis of DHS Data From 64 LMICs
ABSTRACT Background Machine learning (ML) methods are increasingly used for health‐related prediction. This study compared logistic regression with five ML algorithms for predicting reported antibiotic treatment following recent fever/cough among children under five using pooled Demographic and Health Survey (DHS) data. Methods Data from 64 low‐ and middle‐income countries (LMICs) were analyzed. The final analytic sample included 183,425 children with recent fever/cough, of whom 56,769 (30.95%) were reported to have received antibiotics. Logistic regression, K ‐nearest neighbors ( K NN), decision tree, random forest (RF), gradient boosting, and XGBoost were evaluated. The data were divided into model‐development and common heldout test samples, with 55,028 observations in the test set. Hyperparameters were selected using the same three stratified cross‐validation folds for all models, and Synthetic Minority Over‐sampling Technique for Nominal and Continuous (SMOTENC) was applied only within training folds. Model performance was assessed using ROC‐AUC, area under the precision–recall curve (AUPRC), balanced accuracy, sensitivity, specificity, precision, F 1 score, Matthews correlation coefficient (MCC), Brier score, and calibration measures. Primary classification results used a prespecified probability threshold of 0.50. Results Overall predictive performance was modest. RF achieved the highest ROC‐AUC (0.568, 95% CI: 0.552–0.591), whereas K NN achieved the highest AUPRC (0.352, 95% CI: 0.308–0.401) and balanced accuracy (0.531, 95% CI: 0.526–0.536). At the 0.50 threshold, sensitivity was low across all models, ranging from 0.105 to 0.252, whereas specificity ranged from 0.810 to 0.910. Calibration was also limited, with calibration slopes ranging from 0.021 to 0.328. No model consistently performed best across all evaluation measures. Conclusions ML models did not demonstrate clear overall superiority over logistic regression. Although RF and K NN performed better on selected measures, discrimination and calibration remained limited. ML may be useful as a complementary predictive approach, but stronger validation, additional predictors, and more complete consideration of the DHS complex survey design are needed before practical public‐health application.
Authors
- Md Jamal Uddin (ORCID: https://orcid.org/0000-0002-8360-3274)
- Prosenjit Basak Arka (ORCID: https://orcid.org/0009-0000-6159-4883)
- Md. Sabbir Hossain (ORCID: https://orcid.org/0000-0001-7216-0925)
- Mahfuzer Rohman (ORCID: https://orcid.org/0000-0002-1901-2533)
- Md Fakrul Islam
Institutions
- Shahjalal University of Science and Technology (BD)
- Daffodil International University (BD)
Publication Details
- Journal
- Public Health Challenges
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1002/puh2.70413
- Primary Topic
- Antibiotic Use and Resistance
- Type
- article
- Field-Weighted Citation Impact
- 0.00