Modelling a dataset-defined skill-retention outcome in generative AI-assisted learning: An exploratory machine-learning and explainability study

Generative artificial intelligence (GenAI) is increasingly used in higher education, but durable learning cannot be inferred from task completion alone. This study presents an exploratory machine-learning benchmark using a public dataset of 50,000 student-like records that is treated here as synthetic/engineered because its source does not document an empirical sampling frame, institution, country, recruitment process, response rate, or primary-study ethics procedures. The target field, Skill Retention Score, is defined in the source schema on a 0-100 scale as representing skills retained and applied after the semester; however, the source provides no assessment instrument, delayed-assessment interval, reliability estimate, or validity evidence. Accordingly, the variable is analysed as a dataset-defined proxy rather than a validated measure of long-term retention. Ridge Regression, Decision Tree Regression, and Extra Trees Regression were compared in the originally reported 80:20 hold-out analysis. On that split, Extra Trees produced R² = .185, RMSE = 11.983, and MAE = 9.696. Its RMSE is approximately 9.8% lower than the full-sample outcome standard deviation of 13.282, indicating modest predictive gain. The explainability output is permutation importance aggregated to the parent-variable level; it is interpreted as a model-internal, non-directional diagnostic, and SHAP results were not available in the submitted analytical record. The results therefore describe the structure of this benchmark dataset and the behaviour of the modelling pipeline; they do not establish effects of GenAI on real students, directional relationships, or instructional recommendations.

Authors

Institutions

Publication Details

Journal
Journal of Educational Technology and Online Learning
Published
2026-09-30
DOI
https://doi.org/10.31681/jetol.1991829
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Modelling a dataset-defined skill-retention outcome in generative AI-assisted learning: An exploratory machine-learning and explainability study

Bernard Kyiewu, Clinton Amponsah, Linda Bessa-Simons
Journal of Educational Technology and Online Learning
Artificial Intelligence in Healthcare and Education
article

Modelling a dataset-defined skill-retention outcome in generative AI-assisted learning: An exploratory machine-learning and explainability study

Bernard Kyiewu, Clinton Amponsah, Linda Bessa-Simons
article en

Abstract

Generative artificial intelligence (GenAI) is increasingly used in higher education, but durable learning cannot be inferred from task completion alone. This study presents an exploratory machine-learning benchmark using a public dataset of 50,000 student-like records that is treated here as synthetic/engineered because its source does not document an empirical sampling frame, institution, country, recruitment process, response rate, or primary-study ethics procedures. The target field, Skill Retention Score, is defined in the source schema on a 0-100 scale as representing skills retained and applied after the semester; however, the source provides no assessment instrument, delayed-assessment interval, reliability estimate, or validity evidence. Accordingly, the variable is analysed as a dataset-defined proxy rather than a validated measure of long-term retention. Ridge Regression, Decision Tree Regression, and Extra Trees Regression were compared in the originally reported 80:20 hold-out analysis. On that split, Extra Trees produced R² = .185, RMSE = 11.983, and MAE = 9.696. Its RMSE is approximately 9.8% lower than the full-sample outcome standard deviation of 13.282, indicating modest predictive gain. The explainability output is permutation importance aggregated to the parent-variable level; it is interpreted as a model-internal, non-directional diagnostic, and SHAP results were not available in the submitted analytical record. The results therefore describe the structure of this benchmark dataset and the behaviour of the modelling pipeline; they do not establish effects of GenAI on real students, directional relationships, or instructional recommendations.

Journal of Educational Technology and Online LearningVol. 9(3)
University of Energy and Natural Resources (GH)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Modelling a dataset-defined skill-retention outcome in generative AI-assisted learning: An exploratory machine-learning and explainability study — Bernard Kyiewu, Clinton Amponsah, et al. · Journal of Educational Technology and Online Learning (2026) | TGRS Research Map | TGRS