Classification of Educational Vulnerability Status Using Boosting Methods on Imbalanced Socioeconomic Data

Extreme class imbalance in educational data causes classification models to favor the majority class and fail to identify vulnerable students who are not attending school, despite this group requiring the greatest intervention. This study aims to develop a classification model for identifying educational vulnerability status among school-age children. The positive class corresponds to out-of-school children (minority class), while the negative class represents children who are still attending school. The dataset exhibits an extreme imbalance ratio of approximately 1:48, making conventional classification approaches ineffective in detecting minority-class observations.To address this challenge, four imbalance-aware boosting methods, namely AdaBoost-M2, SMOTEBoost, RBBoost, and RUSBoost, were compared. Model selection was conducted using stratified five-fold cross-validation on the training set, followed by final evaluation on an independent test set. Model performance was assessed using sensitivity, specificity, balanced accuracy, and Area Under the Curve (AUC), which are more appropriate than overall accuracy for highly imbalanced data. The results show that SMOTEBoost L50 with a threshold of 0.50 achieved the highest Balanced Accuracy (0.7474), Sensitivity (0.8824), and AUC (0.7745). However, its low minority-class precision (0.0452) and F1-score (0.0860) indicate a substantial false-positive trade-off. These findings demonstrate that integrating minority-class balancing strategies with boosting mechanisms can substantially improve the detection of vulnerable students in highly imbalanced socioeconomic data. Feature importance analysis revealed that education level or class was the most dominant predictor, followed by household size and frequency of internet use. Overall, this study provides empirical evidence regarding the effectiveness of imbalance-aware boosting methods for educational vulnerability detection and highlights their potential application in supporting more targeted educational interventions.

Authors

Institutions

Publication Details

Journal
CAUCHY Jurnal Matematika Murni dan Aplikasi
Published
2026-09-28
DOI
https://doi.org/10.18860/cauchy.v11i2.39670
Primary Topic
Imbalanced Data Classification Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Classification of Educational Vulnerability Status Using Boosting Methods on Imbalanced Socioeconomic Data

Gerry Alfa Dito, Aulia Rizki Firdawanti, Bagus Sartono, Andi Illa Erviani Nensi et al.
CAUCHY Jurnal Matematika Murni dan Aplikasi
Imbalanced Data Classification Techniques
article

Classification of Educational Vulnerability Status Using Boosting Methods on Imbalanced Socioeconomic Data

Gerry Alfa Dito, Aulia Rizki Firdawanti, Bagus Sartono, Andi Illa Erviani Nensi, A. Qeis Tenridapi, Budi Susetyo, Meavi Cintani, Mahda Al Maida
article en

Abstract

Extreme class imbalance in educational data causes classification models to favor the majority class and fail to identify vulnerable students who are not attending school, despite this group requiring the greatest intervention. This study aims to develop a classification model for identifying educational vulnerability status among school-age children. The positive class corresponds to out-of-school children (minority class), while the negative class represents children who are still attending school. The dataset exhibits an extreme imbalance ratio of approximately 1:48, making conventional classification approaches ineffective in detecting minority-class observations.To address this challenge, four imbalance-aware boosting methods, namely AdaBoost-M2, SMOTEBoost, RBBoost, and RUSBoost, were compared. Model selection was conducted using stratified five-fold cross-validation on the training set, followed by final evaluation on an independent test set. Model performance was assessed using sensitivity, specificity, balanced accuracy, and Area Under the Curve (AUC), which are more appropriate than overall accuracy for highly imbalanced data. The results show that SMOTEBoost L50 with a threshold of 0.50 achieved the highest Balanced Accuracy (0.7474), Sensitivity (0.8824), and AUC (0.7745). However, its low minority-class precision (0.0452) and F1-score (0.0860) indicate a substantial false-positive trade-off. These findings demonstrate that integrating minority-class balancing strategies with boosting mechanisms can substantially improve the detection of vulnerable students in highly imbalanced socioeconomic data. Feature importance analysis revealed that education level or class was the most dominant predictor, followed by household size and frequency of internet use. Overall, this study provides empirical evidence regarding the effectiveness of imbalance-aware boosting methods for educational vulnerability detection and highlights their potential application in supporting more targeted educational interventions.

CAUCHY Jurnal Matematika Murni dan AplikasiVol. 11(2)
IPB University (ID)
Openalex Percentile: Top 10%
Imbalanced Data Classification Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.