Tree-Based Machine Learning for Diagnostic Classification of Dengue Fever Using Routine Hematological Parameters: A Secondary Analysis of a Publicly Available Dataset
Background: Dengue fever remains a major global health problem, and early diagnosis is challenging where confirmatory testing is limited. Machine-learning studies using routine hematological data have focused mainly on discrimination, whereas calibration, decision-analytic performance, interpretability, and robust validation have received less attention. This study aimed to develop and compare tree-based machine-learning models for dengue classification, benchmark them against L2-penalized logistic regression (LR), and evaluate discrimination, calibration, potential decision-analytic benefit, and interpretability. Methods: This retrospective secondary analysis used an open-access dataset from Bangladesh comprising 1523 patients, 18 demographic and hematological predictors, and a binary dengue test outcome. Data were divided into stratified training (80%) and test (20%) sets. The Synthetic Minority Over-sampling Technique was applied only within the training workflow. Random Forest (RF), XGBoost, and LightGBM were optimized using Optuna with stratified five-fold cross-validation. L2-penalized LR was evaluated using the same predictors and training–test partition. Held-out test-set performance was assessed using AUROC, AUPRC, accuracy, sensitivity, specificity, predictive values, F1-score, and Brier score. Calibration, decision curve analysis, SHAP values, and permutation importance were also examined. Results: LightGBM, RF, and XGBoost yielded AUROCs of 0.709, 0.704, and 0.702, respectively, indicating closely similar discrimination. The primary SMOTE-trained LR yielded a numerically lower AUROC of 0.608 (95% CI: 0.536–0.677) and a higher Brier score of 0.310 than the tree-based models (0.174–0.176); however, in sensitivity analysis without SMOTE, the LR AUROC increased numerically to 0.655 and the Brier score decreased to 0.194. At the training-derived threshold of 0.558, LightGBM achieved a sensitivity of 0.914 and a specificity of 0.458, reflecting a high-sensitivity, low-specificity profile. The LightGBM calibration curve suggested closer agreement in the low-to-moderate predicted-probability range, with greater deviation at higher probabilities. Decision curve analysis suggested potential net benefit across a range of threshold probabilities but did not establish clinical utility. Platelet count, monocyte percentage, and neutrophil percentage were consistently among the leading predictors across the tree-based models. Conclusions: Tree-based models showed moderate discrimination, with high sensitivity but limited specificity, and yielded numerically higher AUROCs and lower Brier scores than the primary SMOTE-trained LR within this internal-validation framework. They should not replace etiological testing or be used as standalone diagnostic tools; their observed operating characteristics are more compatible with a potential adjunctive screening or triage-support role. External and prospective validation across independent populations and settings is required before clinical use or superiority over simpler statistical models can be established.
Authors
- Sami Akbulut (ORCID: https://orcid.org/0000-0002-6864-7711)
- Zeynep Küçükakçalı (ORCID: https://orcid.org/0000-0001-7956-9272)
- Lecturer Zeynep Burçin Yılmaz (ORCID: https://orcid.org/0000-0002-6950-6013)
Institutions
- Inonu University (TR)
Publication Details
- Journal
- Diagnostics
- Published
- 2026-09-14
- DOI
- https://doi.org/10.3390/diagnostics16182966
- Primary Topic
- Mosquito-borne diseases and control
- Type
- article
- Field-Weighted Citation Impact
- 0.00