Systematic benchmarking of oversampling, undersampling, and hybrid sampling techniques using SMOTE-CCU+ with statistical validation and ablation analysis
Abstract The persistent challenge of class imbalance continues to hinder predictive accuracy and fairness in Educational Data Mining (EDM), particularly in modelling academic performance and student behaviour. In this study, we present a comprehensive benchmark of resampling strategies encompassing oversampling, undersampling, and hybrid approaches to evaluate their effectiveness in improving classification performance across diverse educational contexts. We trained and evaluated five machine learning classifiers (Logistic Regression, Decision Tree, Random Forest, K-Nearest Neighbors, and Support Vector Classifier) on two complementary datasets: a structured academic performance dataset and a publicly available heterogeneous behavioral dataset. To address imbalance-induced bias, we propose a novel hybrid resampling technique, SMOTE-CCU + (Synthetic Minority Oversampling and Cluster Centroid Undersampling Plus), which integrates DBSCAN-based minority filtering, distance-weighted synthetic sample generation, adaptive centroid undersampling (CCU), and Tomek Link boundary refinement. This multi-stage pipeline enhances minority representation while preserving decision-boundary integrity, thereby improving model generalisation. Our experimental results demonstrate that SMOTE-CCU + consistently outperformed established methods (SMOTE, ADASYN, ENN, and CCU) across all evaluation metrics, particularly in the Matthews Correlation Coefficient (MCC), the most reliable indicator of discriminative performance under imbalance. Notably, while tree-based models achieved near-perfect performance even without resampling due to their inherent robustness to class imbalance, SVC’s MCC improved dramatically from 0.000 to 0.986 when paired with SMOTE-CCU + , demonstrating its critical role for classifiers sensitive to imbalance. To ensure rigorous statistical validation, we applied three complementary frameworks: the Wilcoxon Signed-Rank Test with effect size estimation ( $$r = Z/\sqrt{N}$$ ), the Friedman test for multi-group comparison, and Nemenyi post-hoc analysis with critical difference diagrams. The Friedman test confirmed statistically significant differences among sampling strategies on both datasets (primary: $$\chi ^2 = 11.50$$ , $$p = 0.042$$ ; secondary: $$\chi ^2 = 14.07$$ , $$p = 0.015$$ ), while 24 out of 25 pairwise Wilcoxon comparisons yielded large effect sizes ( $$r \ge 0.50$$ ), confirming practical superiority independent of sample size constraints. Our ablation study revealed that DBSCAN filtering is the critical differentiator of the pipeline, while CCU and Tomek Links provide structural refinement most pronounced under noisy, heterogeneous conditions. Computational time analysis further confirmed that SMOTE-CCU + is the most efficient resampling method evaluated (0.036s on the secondary dataset training partition), faster than SMOTE, ENN, ADASYN, and CCU, confirming its suitability for practical deployment. By systematically evaluating our proposed hybrid method alongside standard resampling techniques on both datasets, supported by multi-group statistical validation, effect size reporting, component-level ablation analysis, and computational cost comparison, we provide one of the most comprehensive benchmarks of imbalanced learning in EDM to date. Our findings demonstrate that SMOTE-CCU + is a robust, computationally efficient, and generalisable solution for imbalanced educational data, particularly valuable for institutions deploying diverse classifier portfolios where resampling must benefit all model architectures simultaneously. Through this work, we establish a reproducible and statistically validated benchmark for imbalanced learning in EDM, offering a concrete pathway toward fairer and more reliable predictive systems for student performance monitoring and early intervention.
Authors
- Manal A. Othman (ORCID: https://orcid.org/0000-0002-6246-0761)
- rohayanti binti hassan (ORCID: https://orcid.org/0000-0003-1062-1719)
- Md Abdus Samad (ORCID: https://orcid.org/0000-0002-1990-6924)
- Khaled Mahmud Sujon (ORCID: https://orcid.org/0009-0009-4065-9874)
- Kwonhue Choi
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1038/s41598-026-72493-z
- Primary Topic
- Imbalanced Data Classification Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00