A Comparative Performance Evaluation of Classification Algorithms on Imbalanced Datasets

Class imbalance remains a critical challenge in supervised learning, often biasing classifiers toward majority classes. While resampling techniques like Synthetic Minority Oversampling Technique (SMOTE) are widely used, the combined effect of data balancing and hyperparameter optimization across diverse datasets is rarely systematically explored. This study presents a comprehensive comparative analysis of four classification algorithms—Naive Bayes (NB), K-Nearest Neighbors (K-NN), Artificial Neural Networks (ANN), and Random Forest (RF)—across ten benchmark datasets from the UCI Machine Learning Repository. Unlike previous studies relying on default parameters, this research employs a rigorous Grid Search strategy to optimize hyperparameters for each algorithm within a rigorous SMOTE-balanced stratified cross-validation pipeline to ensure robust evaluation. Performance was assessed using a wide range of metrics, including Accuracy, Precision, Recall, F1-score, and Area Under the Curve (AUC). Experimental results reveal that ANN achieved the highest robustness in high-dimensional and complex categorical datasets (e.g., Car Evaluation F1-score: 0.990), significantly outperforming traditional models. Conversely, RF demonstrated superior stability in datasets with high feature dimensionality (e.g., Arrhythmia F1-score: 0.600) and chemical interactions (e.g., QSAR Fish Toxicity F1-score: 0.839). While K-NN remained competitive in low-dimensional spaces, NB struggled with complex feature dependencies. This study contributes to the literature by demonstrating that algorithmic superiority is context-dependent and providing a data-driven framework for selecting classifiers based on structural characteristics such as dimensionality, categorical complexity, and sample size.

Authors

Institutions

Publication Details

Journal
Sakarya University Journal of Computer and Information Sciences
Published
2026-09-30
DOI
https://doi.org/10.35377/saucis...1833058
Primary Topic
Imbalanced Data Classification Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Comparative Performance Evaluation of Classification Algorithms on Imbalanced Datasets

Necati Vardar, Mehmet Fatih ÖREN
Sakarya University Journal of Computer and Information Sciences
Imbalanced Data Classification Techniques
article

A Comparative Performance Evaluation of Classification Algorithms on Imbalanced Datasets

Necati Vardar, Mehmet Fatih ÖREN
article en

Abstract

Class imbalance remains a critical challenge in supervised learning, often biasing classifiers toward majority classes. While resampling techniques like Synthetic Minority Oversampling Technique (SMOTE) are widely used, the combined effect of data balancing and hyperparameter optimization across diverse datasets is rarely systematically explored. This study presents a comprehensive comparative analysis of four classification algorithms—Naive Bayes (NB), K-Nearest Neighbors (K-NN), Artificial Neural Networks (ANN), and Random Forest (RF)—across ten benchmark datasets from the UCI Machine Learning Repository. Unlike previous studies relying on default parameters, this research employs a rigorous Grid Search strategy to optimize hyperparameters for each algorithm within a rigorous SMOTE-balanced stratified cross-validation pipeline to ensure robust evaluation. Performance was assessed using a wide range of metrics, including Accuracy, Precision, Recall, F1-score, and Area Under the Curve (AUC). Experimental results reveal that ANN achieved the highest robustness in high-dimensional and complex categorical datasets (e.g., Car Evaluation F1-score: 0.990), significantly outperforming traditional models. Conversely, RF demonstrated superior stability in datasets with high feature dimensionality (e.g., Arrhythmia F1-score: 0.600) and chemical interactions (e.g., QSAR Fish Toxicity F1-score: 0.839). While K-NN remained competitive in low-dimensional spaces, NB struggled with complex feature dependencies. This study contributes to the literature by demonstrating that algorithmic superiority is context-dependent and providing a data-driven framework for selecting classifiers based on structural characteristics such as dimensionality, categorical complexity, and sample size.

Sakarya University Journal of Computer and Information SciencesVol. 9(4)
KTO Karatay University (TR)
Life in Land
Openalex Percentile: Top 9%
Imbalanced Data Classification Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A Comparative Performance Evaluation of Classification Algorithms on Imbalanced Datasets — Necati Vardar, Mehmet Fatih ÖREN · Sakarya University Journal of Computer and Information Sciences (2026) | TGRS Research Map | TGRS