RECAP-RL: Symmetry-Guided Retention Shaping and Mirrored Credit Assignment for Personalized Instrumental-Practice Recommendation

Personalized instrumental-practice planning requires selecting and ordering exercises, allocating duration and difficulty, and accounting for prerequisites, fatigue, spacing, and forgetting across multiple sessions. This paper presents RECAP-RL, a symmetry-guided reinforcement-learning framework that post-trains a structured large language model (LLM) policy to generate budget-feasible, teacher-editable practice plans optimized for delayed retention. Symmetry enters at two levels. At the objective level, retention-anchored dense analytic reward shaping (RADAR) re-expresses terminal retention-adjusted learning gain (RALG) as telescoping differences in a retention potential based on the learner’s projected retained mastery; this re-expression preserves core-objective policy ordering and yields an action-independent retention-state baseline whose variance effect is characterized analytically. At the intervention level, matched-intervention replay with residualized order-graph rewards (MIRROR) pairs each exercise with an equal-duration null-practice twin under a common continuation; exchanging the twins reverses the signed retained and prerequisite-unlock contrast, and zero-mean residualization preserves the mean exercise-block advantage when credit is assigned to exercise-token blocks. We instantiate the framework for piano practice in a mechanistic environment with 72 skills and 520 exercises. Against eleven comparison planners on held-out and structurally shifted learner families, RECAP-RL reaches a RALG of 0.399±0.004, exceeding the strongest observable-state model-based planner by 0.034 and terminal-reward group-relative policy optimization by 0.027, while reducing mean cross-family degradation from 26.6% to 18.8% and improving prerequisite repair and spacing alignment. These simulation results support symmetry-guided reward and credit design for delayed educational planning; longitudinal human studies remain necessary to establish educational effectiveness.

Authors

Institutions

Publication Details

Journal
Symmetry
Published
2026-09-29
DOI
https://doi.org/10.3390/sym18101632
Primary Topic
Music Technology and Sound Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

RECAP-RL: Symmetry-Guided Retention Shaping and Mirrored Credit Assignment for Personalized Instrumental-Practice Recommendation

Zhaoen Qu, Zhuodong Liu, Yanlu Li
Symmetry
Music Technology and Sound Studies
article

RECAP-RL: Symmetry-Guided Retention Shaping and Mirrored Credit Assignment for Personalized Instrumental-Practice Recommendation

Zhaoen Qu, Zhuodong Liu, Yanlu Li
article en

Abstract

Personalized instrumental-practice planning requires selecting and ordering exercises, allocating duration and difficulty, and accounting for prerequisites, fatigue, spacing, and forgetting across multiple sessions. This paper presents RECAP-RL, a symmetry-guided reinforcement-learning framework that post-trains a structured large language model (LLM) policy to generate budget-feasible, teacher-editable practice plans optimized for delayed retention. Symmetry enters at two levels. At the objective level, retention-anchored dense analytic reward shaping (RADAR) re-expresses terminal retention-adjusted learning gain (RALG) as telescoping differences in a retention potential based on the learner’s projected retained mastery; this re-expression preserves core-objective policy ordering and yields an action-independent retention-state baseline whose variance effect is characterized analytically. At the intervention level, matched-intervention replay with residualized order-graph rewards (MIRROR) pairs each exercise with an equal-duration null-practice twin under a common continuation; exchanging the twins reverses the signed retained and prerequisite-unlock contrast, and zero-mean residualization preserves the mean exercise-block advantage when credit is assigned to exercise-token blocks. We instantiate the framework for piano practice in a mechanistic environment with 72 skills and 520 exercises. Against eleven comparison planners on held-out and structurally shifted learner families, RECAP-RL reaches a RALG of 0.399±0.004, exceeding the strongest observable-state model-based planner by 0.034 and terminal-reward group-relative policy optimization by 0.027, while reducing mean cross-family degradation from 26.6% to 18.8% and improving prerequisite repair and spacing alignment. These simulation results support symmetry-guided reward and credit design for delayed educational planning; longitudinal human studies remain necessary to establish educational effectiveness.

SymmetryVol. 18(10)
Beijing Jiaotong University (CN), Guangzhou University (CN), University of Hong Kong (HK), South China University of Technology (CN)
Quality Education
Openalex Percentile: Top 14%
Music Technology and Sound Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.