Reward function compression facilitates goal-dependent reinforcement learning
Abstract Humans can uniquely assign value to novel, abstract outcomes to support reinforcement learning. However, this flexibility is cognitively costly and reduces learning efficiency. We propose that goal-dependent learning initially relies on capacity-limited working memory. With consistent experience, learners create a compressed reward function — a simplified rule — that transfers to long-term memory for automatic evaluation upon receiving feedback. This automaticity frees working memory resources, thereby boosting learning efficiency. Across six experiments, we demonstrate that learning is impaired by the size of the goal space but improves when this space allows for compression. Additionally, faster reward processing correlates with better learning. Although the algorithmic details remain to be established, computational modeling revealed that individual differences in compression efficiency, proxied by reward processing speed, led to higher choice accuracy. Together, our behavioral results and computational models suggest that efficient goal-directed learning relies on compressing complex goals into stable reward functions.
Authors
- Gaia Molinaro (ORCID: https://orcid.org/0000-0001-6145-133X)
- Anne Gabrielle Eva Collins (ORCID: https://orcid.org/0000-0003-3751-3662)
Publication Details
- Journal
- Nature Communications
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1038/s41467-026-78295-1
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- article
- Field-Weighted Citation Impact
- 0.00