Comparison of Machine Learning Classifiers for Predicting Reported Antibiotic Treatment After Recent Fever/Cough Among Under‐Five Children: A Pooled Analysis of DHS Data From 64 LMICs

ABSTRACT Background Machine learning (ML) methods are increasingly used for health‐related prediction. This study compared logistic regression with five ML algorithms for predicting reported antibiotic treatment following recent fever/cough among children under five using pooled Demographic and Health Survey (DHS) data. Methods Data from 64 low‐ and middle‐income countries (LMICs) were analyzed. The final analytic sample included 183,425 children with recent fever/cough, of whom 56,769 (30.95%) were reported to have received antibiotics. Logistic regression, K ‐nearest neighbors ( K NN), decision tree, random forest (RF), gradient boosting, and XGBoost were evaluated. The data were divided into model‐development and common heldout test samples, with 55,028 observations in the test set. Hyperparameters were selected using the same three stratified cross‐validation folds for all models, and Synthetic Minority Over‐sampling Technique for Nominal and Continuous (SMOTENC) was applied only within training folds. Model performance was assessed using ROC‐AUC, area under the precision–recall curve (AUPRC), balanced accuracy, sensitivity, specificity, precision, F 1 score, Matthews correlation coefficient (MCC), Brier score, and calibration measures. Primary classification results used a prespecified probability threshold of 0.50. Results Overall predictive performance was modest. RF achieved the highest ROC‐AUC (0.568, 95% CI: 0.552–0.591), whereas K NN achieved the highest AUPRC (0.352, 95% CI: 0.308–0.401) and balanced accuracy (0.531, 95% CI: 0.526–0.536). At the 0.50 threshold, sensitivity was low across all models, ranging from 0.105 to 0.252, whereas specificity ranged from 0.810 to 0.910. Calibration was also limited, with calibration slopes ranging from 0.021 to 0.328. No model consistently performed best across all evaluation measures. Conclusions ML models did not demonstrate clear overall superiority over logistic regression. Although RF and K NN performed better on selected measures, discrimination and calibration remained limited. ML may be useful as a complementary predictive approach, but stronger validation, additional predictors, and more complete consideration of the DHS complex survey design are needed before practical public‐health application.

Authors

Institutions

Publication Details

Journal
Public Health Challenges
Published
2026-10-09
DOI
https://doi.org/10.1002/puh2.70413
Primary Topic
Antibiotic Use and Resistance
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Comparison of Machine Learning Classifiers for Predicting Reported Antibiotic Treatment After Recent Fever/Cough Among Under‐Five Children: A Pooled Analysis of DHS Data From 64 LMICs

Md Jamal Uddin, Prosenjit Basak Arka, Md. Sabbir Hossain, Mahfuzer Rohman et al.
Public Health Challenges
Antibiotic Use and Resistance
article

Comparison of Machine Learning Classifiers for Predicting Reported Antibiotic Treatment After Recent Fever/Cough Among Under‐Five Children: A Pooled Analysis of DHS Data From 64 LMICs

Md Jamal Uddin, Prosenjit Basak Arka, Md. Sabbir Hossain, Mahfuzer Rohman, Md Fakrul Islam
article en

Abstract

ABSTRACT Background Machine learning (ML) methods are increasingly used for health‐related prediction. This study compared logistic regression with five ML algorithms for predicting reported antibiotic treatment following recent fever/cough among children under five using pooled Demographic and Health Survey (DHS) data. Methods Data from 64 low‐ and middle‐income countries (LMICs) were analyzed. The final analytic sample included 183,425 children with recent fever/cough, of whom 56,769 (30.95%) were reported to have received antibiotics. Logistic regression, K ‐nearest neighbors ( K NN), decision tree, random forest (RF), gradient boosting, and XGBoost were evaluated. The data were divided into model‐development and common heldout test samples, with 55,028 observations in the test set. Hyperparameters were selected using the same three stratified cross‐validation folds for all models, and Synthetic Minority Over‐sampling Technique for Nominal and Continuous (SMOTENC) was applied only within training folds. Model performance was assessed using ROC‐AUC, area under the precision–recall curve (AUPRC), balanced accuracy, sensitivity, specificity, precision, F 1 score, Matthews correlation coefficient (MCC), Brier score, and calibration measures. Primary classification results used a prespecified probability threshold of 0.50. Results Overall predictive performance was modest. RF achieved the highest ROC‐AUC (0.568, 95% CI: 0.552–0.591), whereas K NN achieved the highest AUPRC (0.352, 95% CI: 0.308–0.401) and balanced accuracy (0.531, 95% CI: 0.526–0.536). At the 0.50 threshold, sensitivity was low across all models, ranging from 0.105 to 0.252, whereas specificity ranged from 0.810 to 0.910. Calibration was also limited, with calibration slopes ranging from 0.021 to 0.328. No model consistently performed best across all evaluation measures. Conclusions ML models did not demonstrate clear overall superiority over logistic regression. Although RF and K NN performed better on selected measures, discrimination and calibration remained limited. ML may be useful as a complementary predictive approach, but stronger validation, additional predictors, and more complete consideration of the DHS complex survey design are needed before practical public‐health application.

Public Health ChallengesVol. 5(4)
Shahjalal University of Science and Technology (BD), Daffodil International University (BD)
Openalex Percentile: Top 11%
Antibiotic Use and Resistance
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.