What the Reward Curve Didn't Show: Silent Failures in DeepMimic-Style Imitation on the Unitree H1

We reimplement DeepMimic-style motion imitation from scratch for the full 19-DOF Unitree H1 humanoid in MuJoCo MJX and JAX, and report the failures we met on the way. Many of them were silent: training ran, the reward curve rose, and the robot still did the wrong thing. We describe these failures, grouped into action-space, coordinate-frame, data-pipeline, reward-code, training-loop and simulation-setup errors, together with the evidence that exposed each one. Our main observation is that symptoms which look like reward-design problems in motion imitation are often action-space or data-pipeline bugs, and the aggregate reward cannot tell the two apart. Examples include an action scale that made the reference ankle pose unreachable (fixing it raised walking speed from 0.74 to 1.26 m/s), end-effector and centre-of-mass terms that, computed in the world frame as in the original DeepMimic, lost their signal as soon as the pelvis drifted, a centre-of-mass reward that was too flat to give a learning signal, and a sign typo that turned a penalty into a bonus and taught the robot to hop. After the fixes, the plain DeepMimic reward, with no extra path, velocity-command or symmetry terms, produces a policy that walks the full 14.7-minute reference clip in simulation without a single fall, with at most 24 cm of sideways offset from the reference path. We also report a small ablation of termination penalties and a per-joint diagnostic protocol that found the bugs where reward curves did not. This is a case study, not a new method.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23150475
Primary Topic
Reinforcement Learning in Robotics
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

What the Reward Curve Didn't Show: Silent Failures in DeepMimic-Style Imitation on the Unitree H1

Bhargav Oza
Zenodo (CERN European Organization for Nuclear Research)
Reinforcement Learning in Robotics
preprint

What the Reward Curve Didn't Show: Silent Failures in DeepMimic-Style Imitation on the Unitree H1

Bhargav Oza
preprint en

Abstract

We reimplement DeepMimic-style motion imitation from scratch for the full 19-DOF Unitree H1 humanoid in MuJoCo MJX and JAX, and report the failures we met on the way. Many of them were silent: training ran, the reward curve rose, and the robot still did the wrong thing. We describe these failures, grouped into action-space, coordinate-frame, data-pipeline, reward-code, training-loop and simulation-setup errors, together with the evidence that exposed each one. Our main observation is that symptoms which look like reward-design problems in motion imitation are often action-space or data-pipeline bugs, and the aggregate reward cannot tell the two apart. Examples include an action scale that made the reference ankle pose unreachable (fixing it raised walking speed from 0.74 to 1.26 m/s), end-effector and centre-of-mass terms that, computed in the world frame as in the original DeepMimic, lost their signal as soon as the pelvis drifted, a centre-of-mass reward that was too flat to give a learning signal, and a sign typo that turned a penalty into a bonus and taught the robot to hop. After the fixes, the plain DeepMimic reward, with no extra path, velocity-command or symmetry terms, produces a policy that walks the full 14.7-minute reference clip in simulation without a single fall, with at most 24 cm of sideways offset from the reference path. We also report a small ablation of termination penalties and a per-joint diagnostic protocol that found the bugs where reward curves did not. This is a case study, not a new method.

Zenodo (CERN European Organization for Nuclear Research)
Reinforcement Learning in Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.