Reward Shaping based on Trajectory Quality for offline and hybrid reinforcement learning

Current offline reinforcement learning algorithms typically apply equal learning strength to samples of different quality during training, which can limit the effective utilization of high-quality trajectories. To solve this problem, we propose Reward Shaping based on Trajectory Quality (RSTQ). RSTQ increases the rewards of high-quality samples, encouraging the agent to better learn their characteristics and thereby enhancing overall performance. RSTQ is integrated with TD3BC and CQL, and is systematically evaluated on the D4RL benchmark. Experimental results show that RSTQ significantly improves performance on Gym-MuJoCo tasks, with the average normalized returns of TD3BC and CQL increased by 14.2% and 21.6%, respectively, and performance gains exceeding 60% on several tasks, achieving strong competitive results on the evaluated offline reinforcement learning benchmarks. Moreover, RSTQ effectively extends to online reinforcement learning with offline datasets, where it also achieves strong performance.

Authors

Institutions

Publication Details

Journal
Information Processing & Management
Published
2026-09-17
DOI
https://doi.org/10.1016/j.ipm.2026.105160
Primary Topic
Reinforcement Learning in Robotics
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Reward Shaping based on Trajectory Quality for offline and hybrid reinforcement learning

Shuai Lü, Juan Chen, Tao Zhang, Huangyang Chen et al.
Information Processing & Management
Reinforcement Learning in Robotics
article

Reward Shaping based on Trajectory Quality for offline and hybrid reinforcement learning

Shuai Lü, Juan Chen, Tao Zhang, Huangyang Chen, Genghao Sun
article en

Abstract

Current offline reinforcement learning algorithms typically apply equal learning strength to samples of different quality during training, which can limit the effective utilization of high-quality trajectories. To solve this problem, we propose Reward Shaping based on Trajectory Quality (RSTQ). RSTQ increases the rewards of high-quality samples, encouraging the agent to better learn their characteristics and thereby enhancing overall performance. RSTQ is integrated with TD3BC and CQL, and is systematically evaluated on the D4RL benchmark. Experimental results show that RSTQ significantly improves performance on Gym-MuJoCo tasks, with the average normalized returns of TD3BC and CQL increased by 14.2% and 21.6%, respectively, and performance gains exceeding 60% on several tasks, achieving strong competitive results on the evaluated offline reinforcement learning benchmarks. Moreover, RSTQ effectively extends to online reinforcement learning with offline datasets, where it also achieves strong performance.

Information Processing & ManagementVol. 64(2)
Jilin University (CN)
National Natural Science Foundation of China
Openalex Percentile: Top 9%
Reinforcement Learning in Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Reward Shaping based on Trajectory Quality for offline and hybrid reinforcement learning — Shuai Lü, Juan Chen, et al. · Information Processing & Management (2026) | TGRS Research Map | TGRS