Predicting student success in open and distance learning: a comparative evaluation of machine learning models
Purpose This study compares Logistic Regression, linear Support Vector Machine, Random Forest and Histogram Gradient Boosting for predicting four student outcomes in open and distance learning (ODL) from the first four weeks of study. Design/methodology/approach Using the Open University Learning Analytics Dataset, the analysis included 27,535 registrations from 24,841 students. Day-28 predictors comprised static characteristics, raw virtual learning environment (VLE) activity and engagement-focused behavioral features. Class-weighted models were evaluated using nested stratified group cross-validation, with student records grouped across partitions. The macro-averaged F1 score (macro-F1) was the primary metric. Findings With engagement-focused features, the linear support vector machine yielded the highest macro-F1 (0.406) and 47.6% accuracy, followed closely by Random Forest (0.404) and Histogram Gradient Boosting (0.402). Engineered behavioral features improved macro-F1 by 0.014 over raw VLE activity, with a 95% confidence interval of [0.009, 0.019]. A secondary assessment-informed model reached a macro-F1 of 0.434, whereas excluding sensitive demographic attributes did not clearly reduce performance. Originality/value This research contributes to ODL through a fixed early window, student-grouped validation and theory-aligned behavioral feature engineering. It supports timely learner review while treating VLE traces as partial rather than comprehensive indicators of engagement.
Authors
- Mehmet Fırat (ORCID: https://orcid.org/0000-0001-8707-5918)
Institutions
- Anadolu University (TR)
Publication Details
- Journal
- AAOU Journal/AAOU journal
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1108/aaouj-05-2026-0090
- Primary Topic
- Online Learning and Analytics
- Type
- article
- Field-Weighted Citation Impact
- 0.00