Enhancing Imbalanced Data Classification with Class-Aware Synthetic Minority Over-sampling Technique

Imbalanced data classification is one of the difficult machine learning tasks in which the majority class outnumbers the minority classes. This problem is prevalent in various domains such as medical disease detection, spam/fraud detection, digital marketing, agriculture and telecommunications. Synthetic Minority Over-sampling Technique (SMOTE) is a popular oversampling strategy that has been used in addressing the imbalanced dataset classification. However, SMOTE has limitations such as generating noisy samples, ignoring the underlying distribution of the data and over-sampling specific minority regions to the point of over-fitting. To address these limitations, several variants of SMOTE have been proposed, including ADASYN, Borderline-SMOTE, and Safe-Level-SMOTE. However, while these variants aim to address the limitations of SMOTE, they too have their own set of limitations, and their effectiveness may vary across different datasets and problem domains. Therefore, there is still room for improvement in over-sampling techniques to address the limitations of existing methods and improve their performance in diverse scenarios. As a solution, a novel extension of the SMOTE algorithm named "Class-Aware Synthetic Minority over-sampling Technique" (CA-SMOTE) has been proposed. By identifying and utilizing both minority and majority classes to create a synthetic sample. CA-SMOTE generates synthetic samples that are more diverse and closer to the underlying distribution of the data. With this approach, this technique aims to improve the accuracy and reliability of machine learning models trained on imbalanced data. The performance of CA-SMOTE is evaluated on several imbalanced datasets and compared it with SMOTE as well as other modern over-sampling algorithms. The experimental results prove that CA-SMOTE outperforms existing methods in terms of F1-score and AUC-ROC. Overall, CA-SMOTE can be a promising solution for enhanced imbalanced data classification in various real-world applications, improving the accuracy and reliability of machine learning models.

Authors

Institutions

Publication Details

Journal
WSEAS TRANSACTIONS ON COMPUTER RESEARCH
Published
2026-09-29
DOI
https://doi.org/10.37394/232018.2026.14.56
Primary Topic
Imbalanced Data Classification Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Enhancing Imbalanced Data Classification with Class-Aware Synthetic Minority Over-sampling Technique

Ruturaj Mahajan, Vilabha Patil, Sachin Patil
WSEAS TRANSACTIONS ON COMPUTER RESEARCH
Imbalanced Data Classification Techniques
article

Enhancing Imbalanced Data Classification with Class-Aware Synthetic Minority Over-sampling Technique

Ruturaj Mahajan, Vilabha Patil, Sachin Patil
article en

Abstract

Imbalanced data classification is one of the difficult machine learning tasks in which the majority class outnumbers the minority classes. This problem is prevalent in various domains such as medical disease detection, spam/fraud detection, digital marketing, agriculture and telecommunications. Synthetic Minority Over-sampling Technique (SMOTE) is a popular oversampling strategy that has been used in addressing the imbalanced dataset classification. However, SMOTE has limitations such as generating noisy samples, ignoring the underlying distribution of the data and over-sampling specific minority regions to the point of over-fitting. To address these limitations, several variants of SMOTE have been proposed, including ADASYN, Borderline-SMOTE, and Safe-Level-SMOTE. However, while these variants aim to address the limitations of SMOTE, they too have their own set of limitations, and their effectiveness may vary across different datasets and problem domains. Therefore, there is still room for improvement in over-sampling techniques to address the limitations of existing methods and improve their performance in diverse scenarios. As a solution, a novel extension of the SMOTE algorithm named "Class-Aware Synthetic Minority over-sampling Technique" (CA-SMOTE) has been proposed. By identifying and utilizing both minority and majority classes to create a synthetic sample. CA-SMOTE generates synthetic samples that are more diverse and closer to the underlying distribution of the data. With this approach, this technique aims to improve the accuracy and reliability of machine learning models trained on imbalanced data. The performance of CA-SMOTE is evaluated on several imbalanced datasets and compared it with SMOTE as well as other modern over-sampling algorithms. The experimental results prove that CA-SMOTE outperforms existing methods in terms of F1-score and AUC-ROC. Overall, CA-SMOTE can be a promising solution for enhanced imbalanced data classification in various real-world applications, improving the accuracy and reliability of machine learning models.

WSEAS TRANSACTIONS ON COMPUTER RESEARCHVol. 14
Shivaji University (IN)
Openalex Percentile: Top 9%
Imbalanced Data Classification Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Enhancing Imbalanced Data Classification with Class-Aware Synthetic Minority Over-sampling Technique — Ruturaj Mahajan, Vilabha Patil, et al. · WSEAS TRANSACTIONS ON COMPUTER RESEARCH (2026) | TGRS Research Map | TGRS