What the Reward Curve Didn't Show: Silent Failures in DeepMimic-Style Imitation on the Unitree H1
We reimplement DeepMimic-style motion imitation from scratch for the full 19-DOF Unitree H1 humanoid in MuJoCo MJX and JAX, and report the failures we met on the way. Many of them were silent: training ran, the reward curve rose, and the robot still did the wrong thing. We describe these failures, grouped into action-space, coordinate-frame, data-pipeline, reward-code, training-loop and simulation-setup errors, together with the evidence that exposed each one. Our main observation is that symptoms which look like reward-design problems in motion imitation are often action-space or data-pipeline bugs, and the aggregate reward cannot tell the two apart. Examples include an action scale that made the reference ankle pose unreachable (fixing it raised walking speed from 0.74 to 1.26 m/s), end-effector and centre-of-mass terms that, computed in the world frame as in the original DeepMimic, lost their signal as soon as the pelvis drifted, a centre-of-mass reward that was too flat to give a learning signal, and a sign typo that turned a penalty into a bonus and taught the robot to hop. After the fixes, the plain DeepMimic reward, with no extra path, velocity-command or symmetry terms, produces a policy that walks the full 14.7-minute reference clip in simulation without a single fall, with at most 24 cm of sideways offset from the reference path. We also report a small ablation of termination penalties and a per-joint diagnostic protocol that found the bugs where reward curves did not. This is a case study, not a new method.
Authors
- Bhargav Oza
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23150475
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- preprint