Reward Shaping based on Trajectory Quality for offline and hybrid reinforcement learning
Current offline reinforcement learning algorithms typically apply equal learning strength to samples of different quality during training, which can limit the effective utilization of high-quality trajectories. To solve this problem, we propose Reward Shaping based on Trajectory Quality (RSTQ). RSTQ increases the rewards of high-quality samples, encouraging the agent to better learn their characteristics and thereby enhancing overall performance. RSTQ is integrated with TD3BC and CQL, and is systematically evaluated on the D4RL benchmark. Experimental results show that RSTQ significantly improves performance on Gym-MuJoCo tasks, with the average normalized returns of TD3BC and CQL increased by 14.2% and 21.6%, respectively, and performance gains exceeding 60% on several tasks, achieving strong competitive results on the evaluated offline reinforcement learning benchmarks. Moreover, RSTQ effectively extends to online reinforcement learning with offline datasets, where it also achieves strong performance.
Authors
- Shuai Lü (ORCID: https://orcid.org/0000-0002-8081-4498)
- Juan Chen (ORCID: https://orcid.org/0000-0002-9116-6449)
- Tao Zhang (ORCID: https://orcid.org/0000-0002-7992-7962)
- Huangyang Chen (ORCID: https://orcid.org/0009-0009-2272-0985)
- Genghao Sun (ORCID: https://orcid.org/0009-0003-7234-4355)
Institutions
- Jilin University (CN)
Publication Details
- Journal
- Information Processing & Management
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1016/j.ipm.2026.105160
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Natural Science Foundation of China