Calibrated Prediction of Students’ Next Responses in Digital Mathematics Learning: A Completion-Ordered and Cold-Start-Aware Framework

Digital mathematics platforms support probabilistic prediction of subsequent student responses, but temporal ordering, score comparability and calibration affect the validity of evaluation. We evaluated AdaptiveMath-AI on a deterministic cohort of 385,037 Eedi interactions from 2904 students. Early representations were frozen before supervised fitting; learner states were reconstructed in recorded completion order, with simultaneous events predicted before batch updates. Twenty-two numerical inputs supported a gradient-boosting anchor, regularized item corrections, valid curriculum fallback and separately selected calibration. Six monthly expanding-window evaluations, repeated across ten seeds, covered 143,565 responses. The primary estimand averaged within-month performance differences rather than ranking scores from different fitted windows together. Mean monthly ROC-AUC was 0.761173 for AdaptiveMath-AI and 0.756833 for the 100-iteration comparator; the paired difference was 0.004339 (95% conditional student-cluster interval, 0.003636–0.005131). Mean monthly log loss and Brier score were 0.544721 and 0.184383. Independently calibrated ablations supported the contribution of question correction but not an additional hierarchy benefit. Against a recency-matched question-rate control, question-only residual correction improved ROC-AUC by 0.000628 (0.000304–0.000928). Calibration superiority was not uniform. The contribution is a temporally explicit, reproducible evaluation of item adaptation, not a new foundational architecture. Missing presentation timestamps and the previously studied single-platform source limit prospective, external and educational-utility claims.

Authors

Institutions

Publication Details

Journal
Information
Published
2026-10-09
DOI
https://doi.org/10.3390/info17101003
Primary Topic
Intelligent Tutoring Systems and Adaptive Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Calibrated Prediction of Students’ Next Responses in Digital Mathematics Learning: A Completion-Ordered and Cold-Start-Aware Framework

Taganova Guldana, Zhanar Akhayeva, Alma Zakirova, Bakhyt Nurbekov et al.
Information
Intelligent Tutoring Systems and Adaptive Learning
article

Calibrated Prediction of Students’ Next Responses in Digital Mathematics Learning: A Completion-Ordered and Cold-Start-Aware Framework

Taganova Guldana, Zhanar Akhayeva, Alma Zakirova, Bakhyt Nurbekov, Saniya Nariman, Aidana Alibay, Riza Akhitova, Sagynysh Kalmen
article en

Abstract

Digital mathematics platforms support probabilistic prediction of subsequent student responses, but temporal ordering, score comparability and calibration affect the validity of evaluation. We evaluated AdaptiveMath-AI on a deterministic cohort of 385,037 Eedi interactions from 2904 students. Early representations were frozen before supervised fitting; learner states were reconstructed in recorded completion order, with simultaneous events predicted before batch updates. Twenty-two numerical inputs supported a gradient-boosting anchor, regularized item corrections, valid curriculum fallback and separately selected calibration. Six monthly expanding-window evaluations, repeated across ten seeds, covered 143,565 responses. The primary estimand averaged within-month performance differences rather than ranking scores from different fitted windows together. Mean monthly ROC-AUC was 0.761173 for AdaptiveMath-AI and 0.756833 for the 100-iteration comparator; the paired difference was 0.004339 (95% conditional student-cluster interval, 0.003636–0.005131). Mean monthly log loss and Brier score were 0.544721 and 0.184383. Independently calibrated ablations supported the contribution of question correction but not an additional hierarchy benefit. Against a recency-matched question-rate control, question-only residual correction improved ROC-AUC by 0.000628 (0.000304–0.000928). Calibration superiority was not uniform. The contribution is a temporally explicit, reproducible evaluation of item adaptation, not a new foundational architecture. Missing presentation timestamps and the previously studied single-platform source limit prospective, external and educational-utility claims.

InformationVol. 17(10)
L. N. Gumilyov Eurasian National University (KZ), Astana International University, Astana IT University (KZ), International Engineering and Technological University (KZ)
Openalex Percentile: Top 13%
Intelligent Tutoring Systems and Adaptive Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.