Systematic benchmarking of oversampling, undersampling, and hybrid sampling techniques using SMOTE-CCU+ with statistical validation and ablation analysis

Abstract The persistent challenge of class imbalance continues to hinder predictive accuracy and fairness in Educational Data Mining (EDM), particularly in modelling academic performance and student behaviour. In this study, we present a comprehensive benchmark of resampling strategies encompassing oversampling, undersampling, and hybrid approaches to evaluate their effectiveness in improving classification performance across diverse educational contexts. We trained and evaluated five machine learning classifiers (Logistic Regression, Decision Tree, Random Forest, K-Nearest Neighbors, and Support Vector Classifier) on two complementary datasets: a structured academic performance dataset and a publicly available heterogeneous behavioral dataset. To address imbalance-induced bias, we propose a novel hybrid resampling technique, SMOTE-CCU + (Synthetic Minority Oversampling and Cluster Centroid Undersampling Plus), which integrates DBSCAN-based minority filtering, distance-weighted synthetic sample generation, adaptive centroid undersampling (CCU), and Tomek Link boundary refinement. This multi-stage pipeline enhances minority representation while preserving decision-boundary integrity, thereby improving model generalisation. Our experimental results demonstrate that SMOTE-CCU + consistently outperformed established methods (SMOTE, ADASYN, ENN, and CCU) across all evaluation metrics, particularly in the Matthews Correlation Coefficient (MCC), the most reliable indicator of discriminative performance under imbalance. Notably, while tree-based models achieved near-perfect performance even without resampling due to their inherent robustness to class imbalance, SVC’s MCC improved dramatically from 0.000 to 0.986 when paired with SMOTE-CCU + , demonstrating its critical role for classifiers sensitive to imbalance. To ensure rigorous statistical validation, we applied three complementary frameworks: the Wilcoxon Signed-Rank Test with effect size estimation ( $$r = Z/\sqrt{N}$$ ), the Friedman test for multi-group comparison, and Nemenyi post-hoc analysis with critical difference diagrams. The Friedman test confirmed statistically significant differences among sampling strategies on both datasets (primary: $$\chi ^2 = 11.50$$ , $$p = 0.042$$ ; secondary: $$\chi ^2 = 14.07$$ , $$p = 0.015$$ ), while 24 out of 25 pairwise Wilcoxon comparisons yielded large effect sizes ( $$r \ge 0.50$$ ), confirming practical superiority independent of sample size constraints. Our ablation study revealed that DBSCAN filtering is the critical differentiator of the pipeline, while CCU and Tomek Links provide structural refinement most pronounced under noisy, heterogeneous conditions. Computational time analysis further confirmed that SMOTE-CCU + is the most efficient resampling method evaluated (0.036s on the secondary dataset training partition), faster than SMOTE, ENN, ADASYN, and CCU, confirming its suitability for practical deployment. By systematically evaluating our proposed hybrid method alongside standard resampling techniques on both datasets, supported by multi-group statistical validation, effect size reporting, component-level ablation analysis, and computational cost comparison, we provide one of the most comprehensive benchmarks of imbalanced learning in EDM to date. Our findings demonstrate that SMOTE-CCU + is a robust, computationally efficient, and generalisable solution for imbalanced educational data, particularly valuable for institutions deploying diverse classifier portfolios where resampling must benefit all model architectures simultaneously. Through this work, we establish a reproducible and statistically validated benchmark for imbalanced learning in EDM, offering a concrete pathway toward fairer and more reliable predictive systems for student performance monitoring and early intervention.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-10-08
DOI
https://doi.org/10.1038/s41598-026-72493-z
Primary Topic
Imbalanced Data Classification Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Systematic benchmarking of oversampling, undersampling, and hybrid sampling techniques using SMOTE-CCU+ with statistical validation and ablation analysis

Manal A. Othman, rohayanti binti hassan, Md Abdus Samad, Khaled Mahmud Sujon et al.
Scientific Reports
Imbalanced Data Classification Techniques
article

Systematic benchmarking of oversampling, undersampling, and hybrid sampling techniques using SMOTE-CCU+ with statistical validation and ablation analysis

Manal A. Othman, rohayanti binti hassan, Md Abdus Samad, Khaled Mahmud Sujon, Kwonhue Choi
article en

Abstract

Abstract The persistent challenge of class imbalance continues to hinder predictive accuracy and fairness in Educational Data Mining (EDM), particularly in modelling academic performance and student behaviour. In this study, we present a comprehensive benchmark of resampling strategies encompassing oversampling, undersampling, and hybrid approaches to evaluate their effectiveness in improving classification performance across diverse educational contexts. We trained and evaluated five machine learning classifiers (Logistic Regression, Decision Tree, Random Forest, K-Nearest Neighbors, and Support Vector Classifier) on two complementary datasets: a structured academic performance dataset and a publicly available heterogeneous behavioral dataset. To address imbalance-induced bias, we propose a novel hybrid resampling technique, SMOTE-CCU + (Synthetic Minority Oversampling and Cluster Centroid Undersampling Plus), which integrates DBSCAN-based minority filtering, distance-weighted synthetic sample generation, adaptive centroid undersampling (CCU), and Tomek Link boundary refinement. This multi-stage pipeline enhances minority representation while preserving decision-boundary integrity, thereby improving model generalisation. Our experimental results demonstrate that SMOTE-CCU + consistently outperformed established methods (SMOTE, ADASYN, ENN, and CCU) across all evaluation metrics, particularly in the Matthews Correlation Coefficient (MCC), the most reliable indicator of discriminative performance under imbalance. Notably, while tree-based models achieved near-perfect performance even without resampling due to their inherent robustness to class imbalance, SVC’s MCC improved dramatically from 0.000 to 0.986 when paired with SMOTE-CCU + , demonstrating its critical role for classifiers sensitive to imbalance. To ensure rigorous statistical validation, we applied three complementary frameworks: the Wilcoxon Signed-Rank Test with effect size estimation ( $$r = Z/\sqrt{N}$$ ), the Friedman test for multi-group comparison, and Nemenyi post-hoc analysis with critical difference diagrams. The Friedman test confirmed statistically significant differences among sampling strategies on both datasets (primary: $$\chi ^2 = 11.50$$ , $$p = 0.042$$ ; secondary: $$\chi ^2 = 14.07$$ , $$p = 0.015$$ ), while 24 out of 25 pairwise Wilcoxon comparisons yielded large effect sizes ( $$r \ge 0.50$$ ), confirming practical superiority independent of sample size constraints. Our ablation study revealed that DBSCAN filtering is the critical differentiator of the pipeline, while CCU and Tomek Links provide structural refinement most pronounced under noisy, heterogeneous conditions. Computational time analysis further confirmed that SMOTE-CCU + is the most efficient resampling method evaluated (0.036s on the secondary dataset training partition), faster than SMOTE, ENN, ADASYN, and CCU, confirming its suitability for practical deployment. By systematically evaluating our proposed hybrid method alongside standard resampling techniques on both datasets, supported by multi-group statistical validation, effect size reporting, component-level ablation analysis, and computational cost comparison, we provide one of the most comprehensive benchmarks of imbalanced learning in EDM to date. Our findings demonstrate that SMOTE-CCU + is a robust, computationally efficient, and generalisable solution for imbalanced educational data, particularly valuable for institutions deploying diverse classifier portfolios where resampling must benefit all model architectures simultaneously. Through this work, we establish a reproducible and statistically validated benchmark for imbalanced learning in EDM, offering a concrete pathway toward fairer and more reliable predictive systems for student performance monitoring and early intervention.

Scientific Reports
Openalex Percentile: Top 12%
Imbalanced Data Classification Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.