Retrieval‐Augmented Lesson Generation With Large Language Models for Scenario‐Based Tutor Training

ABSTRACT Background High‐impact tutoring can improve student learning, but scaling tutoring programmes requires large numbers of novice tutors. Developing high‐quality tutor‐training lessons remains a key bottleneck. Objectives This study compared post‐test performance after participants studied human‐created or LLM‐created tutor‐training materials and examined whether immediate explanatory feedback moderated differences by lesson creation source (human‐created vs. LLM‐created). Methods We conducted a randomized eight‐condition between‐subjects experiment. Participants completed a source‐matched pre‐instruction practice assessment, studied a tutor‐training lesson created by the same source, and then completed two fixed‐order post‐lesson assessments. The design compared lesson creation source, manipulated immediate explanatory feedback and counterbalanced the pre‐instruction scenario form. Post‐test performance was the mean proportion correct across the human‐created and LLM‐created post‐lesson assessments. Results and Conclusions Average post‐test performance was higher after the human‐created lesson than after the LLM‐created lesson, but this difference depended on feedback. The human‐created advantage appeared in the no‐feedback condition, whereas no statistically reliable lesson creation source difference was detected when feedback was provided. Feedback had no statistically significant overall main effect. The human‐created lesson condition performed better on the human‐created post‐lesson assessment, while the LLM‐created post‐lesson assessment showed a small descriptive difference favouring the LLM‐created lesson condition. Because pre‐instruction performance differed by lesson creation source and scenario form, and because assessment source was confounded with fixed assessment order, the results do not isolate pure learning gains or establish a general transfer advantage.

Authors

Institutions

Publication Details

Journal
Journal of Computer Assisted Learning
Published
2026-10-09
DOI
https://doi.org/10.1002/jcal.70343
Primary Topic
Intelligent Tutoring Systems and Adaptive Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Retrieval‐Augmented Lesson Generation With Large Language Models for Scenario‐Based Tutor Training

Ashish Gurung, Jionghao Lin, Kenneth R. Koedinger, xianghui meng et al.
Journal of Computer Assisted Learning
Intelligent Tutoring Systems and Adaptive Learning
article

Retrieval‐Augmented Lesson Generation With Large Language Models for Scenario‐Based Tutor Training

Ashish Gurung, Jionghao Lin, Kenneth R. Koedinger, xianghui meng, Sandy Yiyang Zhao
article en

Abstract

ABSTRACT Background High‐impact tutoring can improve student learning, but scaling tutoring programmes requires large numbers of novice tutors. Developing high‐quality tutor‐training lessons remains a key bottleneck. Objectives This study compared post‐test performance after participants studied human‐created or LLM‐created tutor‐training materials and examined whether immediate explanatory feedback moderated differences by lesson creation source (human‐created vs. LLM‐created). Methods We conducted a randomized eight‐condition between‐subjects experiment. Participants completed a source‐matched pre‐instruction practice assessment, studied a tutor‐training lesson created by the same source, and then completed two fixed‐order post‐lesson assessments. The design compared lesson creation source, manipulated immediate explanatory feedback and counterbalanced the pre‐instruction scenario form. Post‐test performance was the mean proportion correct across the human‐created and LLM‐created post‐lesson assessments. Results and Conclusions Average post‐test performance was higher after the human‐created lesson than after the LLM‐created lesson, but this difference depended on feedback. The human‐created advantage appeared in the no‐feedback condition, whereas no statistically reliable lesson creation source difference was detected when feedback was provided. Feedback had no statistically significant overall main effect. The human‐created lesson condition performed better on the human‐created post‐lesson assessment, while the LLM‐created post‐lesson assessment showed a small descriptive difference favouring the LLM‐created lesson condition. Because pre‐instruction performance differed by lesson creation source and scenario form, and because assessment source was confounded with fixed assessment order, the results do not isolate pure learning gains or establish a general transfer advantage.

Journal of Computer Assisted LearningVol. 42(6)
Carnegie Mellon University (US), University of Hong Kong (HK)
Openalex Percentile: Top 12%
Intelligent Tutoring Systems and Adaptive Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.