Early Academic Performance Prediction in Secondary Education: Are Simple Machine Learning Models Enough?
Most predictive approaches in educational data mining rely on complex models whose opacity limits practical adoption by classroom teachers, creating a gap between model sophistication and classroom usability. This gap is particularly acute at the class-group level, where institutional gradebook data are routinely aggregated for teacher-level planning but rarely modelled with an explicit account of when model complexity is actually justified. This paper addresses that gap: its novelty is to provide a structural explanation, grounded in group-level academic dynamics, for why linear models are highly competitive, rather than merely adequate, for this type of data, and to test this account empirically. An eight-year longitudinal dataset (2013/2014–2020/2021) from a Spanish secondary school—1070 class-group records across 32 subjects—was used to compare linear regression and Random Forest for final grade prediction, a Random Forest classifier against an XGBoost classifier for academic risk detection, and SHAP (SHapley Additive exPlanations)-based explainability, validated through Leave-One-Course-Out (LOCO) cross-validation. Within this dataset, linear regression consistently matches or outperforms Random Forest in both scenarios (R2 = 0.857 with two assessments; R2 = 0.740 with one), explained by stable cohort dynamics—baseline grades, teaching continuity, group composition—that produce a linear temporal structure (Spearman ρ > 0.81) leaving little predictive return for ensemble complexity in this setting. For the passing class, the Random Forest classifier achieves F1 = 0.972 with high inter-cohort stability (LOCO F1 ∈ [0.944, 0.984]); for the minority at-risk class, it outperforms XGBoost (F1 = 0.69 vs. 0.57), a gap consistent with the benefit of explicit class-imbalance handling, though fully disentangling this from a possible ensemble-family effect is left for future work. The 2019/2020 cohort is statistically anomalous (Mann–Whitney U, p < 0.001), reflecting an exogenous shift in the grade-generating process under emergency evaluation rather than evidence against the linearity account under normal conditions. Simple, transparent models operating on routinely collected gradebook data deliver actionable early-warning signals within the digital competence of most practising teachers; group-level prediction additionally protects student identity by ensuring no individual is labelled at-risk, combining predictive utility with ethical design.
Authors
- Carlos M. Travieso (ORCID: https://orcid.org/0000-0002-4621-2768)
- Miriam Martín-Paciente (ORCID: https://orcid.org/0000-0003-1431-8266)
- Carmen Román-León (ORCID: https://orcid.org/0009-0009-3616-174X)
- Víctor D. Díaz Suárez (ORCID: https://orcid.org/0000-0002-8146-176X)
- María de los Ángeles Buenavista-Ruiz (ORCID: https://orcid.org/0009-0004-1466-0752)
- Marina Praena-Delgado
Institutions
- Universidad de Las Palmas de Gran Canaria (ES)
- Universidad Internacional De La Rioja (ES)
Publication Details
- Journal
- Applied System Innovation
- Published
- 2026-09-11
- DOI
- https://doi.org/10.3390/asi9090191
- Primary Topic
- Online Learning and Analytics
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Universidad de La Laguna
- Universidad de Las Palmas de Gran Canaria