Evaluating the performance of ensemble learning methods in diabetes disease classification

Abstract Diabetes mellitus is a widespread metabolic disorder marked by chronic hyperglycemia and severe complications. Early and accurate detection is crucial for effective management and preventing disease progression. This study systematically evaluates the performance of three ensemble learning strategies Bagging, Boosting, and Stacking on three benchmark diabetes datasets: Pima Indians Diabetes (PID), Frankfurt Hospital Diabetes, and Sylhet Hospital Diabetes. To address class imbalance while preventing data leakage, the Synthetic Minority Oversampling Technique (SMOTE) was applied exclusively to the training data after train-test splitting and independently within each cross-validation fold. Furthermore, all models were evaluated using stratified k-fold cross-validation, and performance was assessed using accuracy, precision, recall, F1-score, receiver operating characteristic area under the curve (ROC-AUC), and calibration analysis. Statistical significance testing was additionally conducted to compare the performance of competing ensemble methods. Experimental results show that all three ensemble paradigms achieved strong performance after SMOTE, with the best-performing model varying by dataset rather than one paradigm uniformly dominating. On the PID dataset, Light Gradient Boosting achieved the highest accuracy (75.97%) On the Frankfurt dataset Bagging and Light Gradient Boosting reached the highest accuracy (98.50%), while on the Sylhet dataset, Bagging perfect accuracy (99.09%) closely followed by Random Forest, Extra Trees and Gradient Boosting (99.03%). These findings underscore the effectiveness of combining SMOTE with Boosting-based ensembles to mitigate class imbalance and improve diabetes classification, highlighting the critical role of both data preprocessing and algorithm selection in achieving high predictive performance for precision medicine and clinical decision support.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-09-29
DOI
https://doi.org/10.1038/s41598-026-73245-9
Primary Topic
Artificial Intelligence in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating the performance of ensemble learning methods in diabetes disease classification

Sajjad aghasi javid, Aliasghar Khakpaki
Scientific Reports
Artificial Intelligence in Healthcare
article

Evaluating the performance of ensemble learning methods in diabetes disease classification

Sajjad aghasi javid, Aliasghar Khakpaki
article en

Abstract

Abstract Diabetes mellitus is a widespread metabolic disorder marked by chronic hyperglycemia and severe complications. Early and accurate detection is crucial for effective management and preventing disease progression. This study systematically evaluates the performance of three ensemble learning strategies Bagging, Boosting, and Stacking on three benchmark diabetes datasets: Pima Indians Diabetes (PID), Frankfurt Hospital Diabetes, and Sylhet Hospital Diabetes. To address class imbalance while preventing data leakage, the Synthetic Minority Oversampling Technique (SMOTE) was applied exclusively to the training data after train-test splitting and independently within each cross-validation fold. Furthermore, all models were evaluated using stratified k-fold cross-validation, and performance was assessed using accuracy, precision, recall, F1-score, receiver operating characteristic area under the curve (ROC-AUC), and calibration analysis. Statistical significance testing was additionally conducted to compare the performance of competing ensemble methods. Experimental results show that all three ensemble paradigms achieved strong performance after SMOTE, with the best-performing model varying by dataset rather than one paradigm uniformly dominating. On the PID dataset, Light Gradient Boosting achieved the highest accuracy (75.97%) On the Frankfurt dataset Bagging and Light Gradient Boosting reached the highest accuracy (98.50%), while on the Sylhet dataset, Bagging perfect accuracy (99.09%) closely followed by Random Forest, Extra Trees and Gradient Boosting (99.03%). These findings underscore the effectiveness of combining SMOTE with Boosting-based ensembles to mitigate class imbalance and improve diabetes classification, highlighting the critical role of both data preprocessing and algorithm selection in achieving high predictive performance for precision medicine and clinical decision support.

Scientific Reports
Good health and well-being
Openalex Percentile: Top 3%
Artificial Intelligence in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.