Performance enhancement of machine learning models for diabetes diagnosis through data balancing and hyperparameter optimization

Early and reliable prediction of diabetes is crucial for better outcomes. Seven machine learning (ML) algorithms—Logistic Regression (LR), Decision Tree (DT), Support Vector Machines (SVM), K-Nearest Neighbors (KNN), Gradient Boosting (GB), LightGBM, and CatBoost—are assessed for diagnosing diabetes automatically. The models were tested using a sequential approach, including imputing missing values, removing outliers, balancing the data samples using Synthetic Minority Oversampling Technique (SMOTE) and hyperparameter tuning using GridSearchCV. Using 10-fold cross-validation on the Pima Indian Diabetes (PIMA) dataset, CatBoost outperformed other models with an accuracy of 83.70% and F1-score of 84.49%, with KNN, GB and LightGBM performing similarly. External validation on a German dataset ( n = 2,000) yielded higher accuracy (up to ~90%), with KNN, LightGBM, and CatBoost maintaining strong performance, suggesting good generalization across datasets. Overall, the results indicate that careful preprocessing and model optimization can enhance predictive performance, highlighting the potential of ensemble-based methods for clinical decision support.

Authors

Institutions

Publication Details

Journal
PeerJ Computer Science
Published
2026-10-06
DOI
https://doi.org/10.7717/peerj-cs.4099
Primary Topic
Artificial Intelligence in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Performance enhancement of machine learning models for diabetes diagnosis through data balancing and hyperparameter optimization

Neelam Gohar, Muhammad Uzair Khan, Amal Al‐Rasheed, Fazal Muhammad et al.
PeerJ Computer Science
Artificial Intelligence in Healthcare
article

Performance enhancement of machine learning models for diabetes diagnosis through data balancing and hyperparameter optimization

Neelam Gohar, Muhammad Uzair Khan, Amal Al‐Rasheed, Fazal Muhammad, Shahid Khan, Saadia Tabassum, Muhammad Lais
article en

Abstract

Early and reliable prediction of diabetes is crucial for better outcomes. Seven machine learning (ML) algorithms—Logistic Regression (LR), Decision Tree (DT), Support Vector Machines (SVM), K-Nearest Neighbors (KNN), Gradient Boosting (GB), LightGBM, and CatBoost—are assessed for diagnosing diabetes automatically. The models were tested using a sequential approach, including imputing missing values, removing outliers, balancing the data samples using Synthetic Minority Oversampling Technique (SMOTE) and hyperparameter tuning using GridSearchCV. Using 10-fold cross-validation on the Pima Indian Diabetes (PIMA) dataset, CatBoost outperformed other models with an accuracy of 83.70% and F1-score of 84.49%, with KNN, GB and LightGBM performing similarly. External validation on a German dataset ( n = 2,000) yielded higher accuracy (up to ~90%), with KNN, LightGBM, and CatBoost maintaining strong performance, suggesting good generalization across datasets. Overall, the results indicate that careful preprocessing and model optimization can enhance predictive performance, highlighting the potential of ensemble-based methods for clinical decision support.

PeerJ Computer ScienceVol. 12
Princess Nourah bint Abdulrahman University (SA), Shaheed Benazir Bhutto Women University Peshawar (PK), COMSATS University Islamabad (PK), Abdul Wali Khan University Mardan (PK), University of Technology Nowshera, University of Engineering and Technology Peshawar (PK)
Openalex Percentile: Top 7%
Artificial Intelligence in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.