Tree-Based Machine Learning for Diagnostic Classification of Dengue Fever Using Routine Hematological Parameters: A Secondary Analysis of a Publicly Available Dataset

Background: Dengue fever remains a major global health problem, and early diagnosis is challenging where confirmatory testing is limited. Machine-learning studies using routine hematological data have focused mainly on discrimination, whereas calibration, decision-analytic performance, interpretability, and robust validation have received less attention. This study aimed to develop and compare tree-based machine-learning models for dengue classification, benchmark them against L2-penalized logistic regression (LR), and evaluate discrimination, calibration, potential decision-analytic benefit, and interpretability. Methods: This retrospective secondary analysis used an open-access dataset from Bangladesh comprising 1523 patients, 18 demographic and hematological predictors, and a binary dengue test outcome. Data were divided into stratified training (80%) and test (20%) sets. The Synthetic Minority Over-sampling Technique was applied only within the training workflow. Random Forest (RF), XGBoost, and LightGBM were optimized using Optuna with stratified five-fold cross-validation. L2-penalized LR was evaluated using the same predictors and training–test partition. Held-out test-set performance was assessed using AUROC, AUPRC, accuracy, sensitivity, specificity, predictive values, F1-score, and Brier score. Calibration, decision curve analysis, SHAP values, and permutation importance were also examined. Results: LightGBM, RF, and XGBoost yielded AUROCs of 0.709, 0.704, and 0.702, respectively, indicating closely similar discrimination. The primary SMOTE-trained LR yielded a numerically lower AUROC of 0.608 (95% CI: 0.536–0.677) and a higher Brier score of 0.310 than the tree-based models (0.174–0.176); however, in sensitivity analysis without SMOTE, the LR AUROC increased numerically to 0.655 and the Brier score decreased to 0.194. At the training-derived threshold of 0.558, LightGBM achieved a sensitivity of 0.914 and a specificity of 0.458, reflecting a high-sensitivity, low-specificity profile. The LightGBM calibration curve suggested closer agreement in the low-to-moderate predicted-probability range, with greater deviation at higher probabilities. Decision curve analysis suggested potential net benefit across a range of threshold probabilities but did not establish clinical utility. Platelet count, monocyte percentage, and neutrophil percentage were consistently among the leading predictors across the tree-based models. Conclusions: Tree-based models showed moderate discrimination, with high sensitivity but limited specificity, and yielded numerically higher AUROCs and lower Brier scores than the primary SMOTE-trained LR within this internal-validation framework. They should not replace etiological testing or be used as standalone diagnostic tools; their observed operating characteristics are more compatible with a potential adjunctive screening or triage-support role. External and prospective validation across independent populations and settings is required before clinical use or superiority over simpler statistical models can be established.

Authors

Institutions

Publication Details

Journal
Diagnostics
Published
2026-09-14
DOI
https://doi.org/10.3390/diagnostics16182966
Primary Topic
Mosquito-borne diseases and control
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Tree-Based Machine Learning for Diagnostic Classification of Dengue Fever Using Routine Hematological Parameters: A Secondary Analysis of a Publicly Available Dataset

Sami Akbulut, Zeynep Küçükakçalı, Lecturer Zeynep Burçin Yılmaz
Diagnostics
Mosquito-borne diseases and control
article

Tree-Based Machine Learning for Diagnostic Classification of Dengue Fever Using Routine Hematological Parameters: A Secondary Analysis of a Publicly Available Dataset

Sami Akbulut, Zeynep Küçükakçalı, Lecturer Zeynep Burçin Yılmaz
article en

Abstract

Background: Dengue fever remains a major global health problem, and early diagnosis is challenging where confirmatory testing is limited. Machine-learning studies using routine hematological data have focused mainly on discrimination, whereas calibration, decision-analytic performance, interpretability, and robust validation have received less attention. This study aimed to develop and compare tree-based machine-learning models for dengue classification, benchmark them against L2-penalized logistic regression (LR), and evaluate discrimination, calibration, potential decision-analytic benefit, and interpretability. Methods: This retrospective secondary analysis used an open-access dataset from Bangladesh comprising 1523 patients, 18 demographic and hematological predictors, and a binary dengue test outcome. Data were divided into stratified training (80%) and test (20%) sets. The Synthetic Minority Over-sampling Technique was applied only within the training workflow. Random Forest (RF), XGBoost, and LightGBM were optimized using Optuna with stratified five-fold cross-validation. L2-penalized LR was evaluated using the same predictors and training–test partition. Held-out test-set performance was assessed using AUROC, AUPRC, accuracy, sensitivity, specificity, predictive values, F1-score, and Brier score. Calibration, decision curve analysis, SHAP values, and permutation importance were also examined. Results: LightGBM, RF, and XGBoost yielded AUROCs of 0.709, 0.704, and 0.702, respectively, indicating closely similar discrimination. The primary SMOTE-trained LR yielded a numerically lower AUROC of 0.608 (95% CI: 0.536–0.677) and a higher Brier score of 0.310 than the tree-based models (0.174–0.176); however, in sensitivity analysis without SMOTE, the LR AUROC increased numerically to 0.655 and the Brier score decreased to 0.194. At the training-derived threshold of 0.558, LightGBM achieved a sensitivity of 0.914 and a specificity of 0.458, reflecting a high-sensitivity, low-specificity profile. The LightGBM calibration curve suggested closer agreement in the low-to-moderate predicted-probability range, with greater deviation at higher probabilities. Decision curve analysis suggested potential net benefit across a range of threshold probabilities but did not establish clinical utility. Platelet count, monocyte percentage, and neutrophil percentage were consistently among the leading predictors across the tree-based models. Conclusions: Tree-based models showed moderate discrimination, with high sensitivity but limited specificity, and yielded numerically higher AUROCs and lower Brier scores than the primary SMOTE-trained LR within this internal-validation framework. They should not replace etiological testing or be used as standalone diagnostic tools; their observed operating characteristics are more compatible with a potential adjunctive screening or triage-support role. External and prospective validation across independent populations and settings is required before clinical use or superiority over simpler statistical models can be established.

DiagnosticsVol. 16(18)
Inonu University (TR)
Peace, Justice and strong institutions, Reduced inequalities
Openalex Percentile: Top 8%
Mosquito-borne diseases and control
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.