Toward Sustainable and Equitable AI in Education Through a Regional Fairness Audit of Dropout Prediction Using the OULAD Dataset

Ensuring that artificial intelligence contributes to sustainable, equitable education requires more than aggregate accuracy—it requires verifying that predictive systems serve all learners fairly, including across geographic regions. We audit a dropout-prediction pipeline built on the Open University Learning Analytics Dataset (OULAD) for disparities across gender, disability, and geographic region, using a student-level train/test partition to prevent the same student’s records from contaminating both sets. Logistic regression and random forest classifiers attain approximately 0.85 accuracy and 0.91–0.92 AUC overall, yet region-stratified recall (true-positive rate) ranges from 0.55 in Wales to 0.80 in the West Midlands Region, an equal-opportunity gap of 0.25 that is corroborated by region-specific AUC, by a random forest classifier, and, for actual withdrawals, by a likelihood-ratio test showing region predicts being missed by the classifier beyond what the Index of Multiple Deprivation (IMD) explains. A region-isolation test shows that excluding region as a model predictor is associated with a significantly narrower gap in both model families (0.13–0.19 without region versus 0.25–0.28 with region), an association not explained by IMD band alone; because the bootstrap 95% confidence interval on this difference ([−0.003,0.173]) narrowly includes zero, we treat the attribution to region specifically as suggestive rather than conclusively established. A naive region-specific decision-threshold mitigation, evaluated correctly on a held-out validation set, does not improve the gap; a shrinkage-regularized version recovers a modest, observed reduction (0.24 to 0.17) on the held-out test set, without a formal uncertainty interval for this difference, at the cost of a near-doubling of regional false-positive rates. Because our analysis is retrospective and several predictors are computed over the full module presentation, these findings support methodological lessons for the design and auditing of future systems rather than direct claims about the performance or fairness of an operational, real-time early-warning intervention; we report them, including the mitigation failure, as evidence that dropout-prediction systems audited only for aggregate accuracy, without a properly validated regional fairness assessment, risk under-serving or unevenly burdening students in specific regions, working against rather than for the aims of SDG 4.

Authors

Institutions

Publication Details

Journal
Sustainability
Published
2026-09-15
DOI
https://doi.org/10.3390/su18189440
Primary Topic
Online Learning and Analytics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Toward Sustainable and Equitable AI in Education Through a Regional Fairness Audit of Dropout Prediction Using the OULAD Dataset

Yousef Wardat, Firuz Kamalov, Ahmed Abdel Aziz Elsayed, Hana Sulieman
Sustainability
Online Learning and Analytics
article

Toward Sustainable and Equitable AI in Education Through a Regional Fairness Audit of Dropout Prediction Using the OULAD Dataset

Yousef Wardat, Firuz Kamalov, Ahmed Abdel Aziz Elsayed, Hana Sulieman
article en

Abstract

Ensuring that artificial intelligence contributes to sustainable, equitable education requires more than aggregate accuracy—it requires verifying that predictive systems serve all learners fairly, including across geographic regions. We audit a dropout-prediction pipeline built on the Open University Learning Analytics Dataset (OULAD) for disparities across gender, disability, and geographic region, using a student-level train/test partition to prevent the same student’s records from contaminating both sets. Logistic regression and random forest classifiers attain approximately 0.85 accuracy and 0.91–0.92 AUC overall, yet region-stratified recall (true-positive rate) ranges from 0.55 in Wales to 0.80 in the West Midlands Region, an equal-opportunity gap of 0.25 that is corroborated by region-specific AUC, by a random forest classifier, and, for actual withdrawals, by a likelihood-ratio test showing region predicts being missed by the classifier beyond what the Index of Multiple Deprivation (IMD) explains. A region-isolation test shows that excluding region as a model predictor is associated with a significantly narrower gap in both model families (0.13–0.19 without region versus 0.25–0.28 with region), an association not explained by IMD band alone; because the bootstrap 95% confidence interval on this difference ([−0.003,0.173]) narrowly includes zero, we treat the attribution to region specifically as suggestive rather than conclusively established. A naive region-specific decision-threshold mitigation, evaluated correctly on a held-out validation set, does not improve the gap; a shrinkage-regularized version recovers a modest, observed reduction (0.24 to 0.17) on the held-out test set, without a formal uncertainty interval for this difference, at the cost of a near-doubling of regional false-positive rates. Because our analysis is retrospective and several predictors are computed over the full module presentation, these findings support methodological lessons for the design and auditing of future systems rather than direct claims about the performance or fairness of an operational, real-time early-warning intervention; we report them, including the mitigation failure, as evidence that dropout-prediction systems audited only for aggregate accuracy, without a properly validated regional fairness assessment, risk under-serving or unevenly burdening students in specific regions, working against rather than for the aims of SDG 4.

SustainabilityVol. 18(18)
Canadian University of Dubai (AE), American University of Sharjah (AE), Yarmouk University (JO)
Quality Education
Openalex Percentile: Top 5%
Online Learning and Analytics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.