Early Academic Performance Prediction in Secondary Education: Are Simple Machine Learning Models Enough?

Most predictive approaches in educational data mining rely on complex models whose opacity limits practical adoption by classroom teachers, creating a gap between model sophistication and classroom usability. This gap is particularly acute at the class-group level, where institutional gradebook data are routinely aggregated for teacher-level planning but rarely modelled with an explicit account of when model complexity is actually justified. This paper addresses that gap: its novelty is to provide a structural explanation, grounded in group-level academic dynamics, for why linear models are highly competitive, rather than merely adequate, for this type of data, and to test this account empirically. An eight-year longitudinal dataset (2013/2014–2020/2021) from a Spanish secondary school—1070 class-group records across 32 subjects—was used to compare linear regression and Random Forest for final grade prediction, a Random Forest classifier against an XGBoost classifier for academic risk detection, and SHAP (SHapley Additive exPlanations)-based explainability, validated through Leave-One-Course-Out (LOCO) cross-validation. Within this dataset, linear regression consistently matches or outperforms Random Forest in both scenarios (R2 = 0.857 with two assessments; R2 = 0.740 with one), explained by stable cohort dynamics—baseline grades, teaching continuity, group composition—that produce a linear temporal structure (Spearman ρ > 0.81) leaving little predictive return for ensemble complexity in this setting. For the passing class, the Random Forest classifier achieves F1 = 0.972 with high inter-cohort stability (LOCO F1 ∈ [0.944, 0.984]); for the minority at-risk class, it outperforms XGBoost (F1 = 0.69 vs. 0.57), a gap consistent with the benefit of explicit class-imbalance handling, though fully disentangling this from a possible ensemble-family effect is left for future work. The 2019/2020 cohort is statistically anomalous (Mann–Whitney U, p < 0.001), reflecting an exogenous shift in the grade-generating process under emergency evaluation rather than evidence against the linearity account under normal conditions. Simple, transparent models operating on routinely collected gradebook data deliver actionable early-warning signals within the digital competence of most practising teachers; group-level prediction additionally protects student identity by ensuring no individual is labelled at-risk, combining predictive utility with ethical design.

Authors

Institutions

Publication Details

Journal
Applied System Innovation
Published
2026-09-11
DOI
https://doi.org/10.3390/asi9090191
Primary Topic
Online Learning and Analytics
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Early Academic Performance Prediction in Secondary Education: Are Simple Machine Learning Models Enough?

Carlos M. Travieso, Miriam Martín-Paciente, Carmen Román-León, Víctor D. Díaz Suárez et al.
Applied System Innovation
Online Learning and Analytics
article

Early Academic Performance Prediction in Secondary Education: Are Simple Machine Learning Models Enough?

Carlos M. Travieso, Miriam Martín-Paciente, Carmen Román-León, Víctor D. Díaz Suárez, María de los Ángeles Buenavista-Ruiz, Marina Praena-Delgado
article en

Abstract

Most predictive approaches in educational data mining rely on complex models whose opacity limits practical adoption by classroom teachers, creating a gap between model sophistication and classroom usability. This gap is particularly acute at the class-group level, where institutional gradebook data are routinely aggregated for teacher-level planning but rarely modelled with an explicit account of when model complexity is actually justified. This paper addresses that gap: its novelty is to provide a structural explanation, grounded in group-level academic dynamics, for why linear models are highly competitive, rather than merely adequate, for this type of data, and to test this account empirically. An eight-year longitudinal dataset (2013/2014–2020/2021) from a Spanish secondary school—1070 class-group records across 32 subjects—was used to compare linear regression and Random Forest for final grade prediction, a Random Forest classifier against an XGBoost classifier for academic risk detection, and SHAP (SHapley Additive exPlanations)-based explainability, validated through Leave-One-Course-Out (LOCO) cross-validation. Within this dataset, linear regression consistently matches or outperforms Random Forest in both scenarios (R2 = 0.857 with two assessments; R2 = 0.740 with one), explained by stable cohort dynamics—baseline grades, teaching continuity, group composition—that produce a linear temporal structure (Spearman ρ > 0.81) leaving little predictive return for ensemble complexity in this setting. For the passing class, the Random Forest classifier achieves F1 = 0.972 with high inter-cohort stability (LOCO F1 ∈ [0.944, 0.984]); for the minority at-risk class, it outperforms XGBoost (F1 = 0.69 vs. 0.57), a gap consistent with the benefit of explicit class-imbalance handling, though fully disentangling this from a possible ensemble-family effect is left for future work. The 2019/2020 cohort is statistically anomalous (Mann–Whitney U, p < 0.001), reflecting an exogenous shift in the grade-generating process under emergency evaluation rather than evidence against the linearity account under normal conditions. Simple, transparent models operating on routinely collected gradebook data deliver actionable early-warning signals within the digital competence of most practising teachers; group-level prediction additionally protects student identity by ensuring no individual is labelled at-risk, combining predictive utility with ethical design.

Applied System InnovationVol. 9(9)
Universidad de Las Palmas de Gran Canaria (ES), Universidad Internacional De La Rioja (ES)
Universidad de La Laguna, Universidad de Las Palmas de Gran Canaria
Quality Education
Openalex Percentile: Top 5%
Online Learning and Analytics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.