Explainable AI for predicting and interpreting mathematics achievement: a cross-national analysis of PISA 2018
Abstract This study applies a survey-weighted, plausible-value-aware explainable machine-learning workflow to PISA 2018 mathematics data from 74,235 students in ten purposively selected education systems. The analysis examines cross-national stability and country-specific variation in prediction and interpretation within a comparative, non-representative sample. A primary pool of 33 student questionnaire predictors and OECD-derived indices was constructed, followed by country-specific train-only stability feature selection. Four weighted models were compared using school-level holdout test sets: weighted linear regression, weighted linear regression with pairwise interactions, random forest, and CatBoost. Model performance was evaluated across all ten mathematics plausible values using final student weights and replicate-weight variance estimation. CatBoost achieved the strongest performance across all systems, with an average weighted $$\\:{R}^{2}$$ of 0.358 and an average MAE of 57.29 score points. PV-robust SHAP summaries showed that books at home, highest parental occupational status, grade placement, mathematics learning time, socioeconomic status, test effort, directed instruction, and disciplinary climate were among the most stable predictors. Sensitivity analyses indicated that the SHAP pattern was not driven by PV1MATH or by grade placement alone. The study contributes a reproducible framework for explainable machine learning in comparative large-scale assessment research.
Authors
- Liu Liu (ORCID: https://orcid.org/0000-0001-6900-9829)
- Rui Dai (ORCID: https://orcid.org/0000-0001-5477-8078)
Publication Details
- Journal
- Large-scale Assessments in Education
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1186/s40536-026-00320-y
- Primary Topic
- Intelligent Tutoring Systems and Adaptive Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00