Decoupling and conditioning reshape influence allocation and the gradient-noise floor in multi-reward GRPO under a finite-sample U-statistic analysis
Abstract Multi-reward reinforcement learning with verifiable rewards (RLVR) increasingly relies on Group Relative Policy Optimization (GRPO). Two recent methods—MO-GRPO and GDPO—replace GRPO’s standard aggregate-then-normalize (AN) advantage with a decoupled normalize-then-aggregate (NA) estimator, and show empirically that AN lets the highest-variance reward channel dominate and collapses distinct reward combinations into identical advantages. We supply the missing finite-sample, correlation-aware theory of these estimators. Our two main results are: (i) an exact finite-sample mean-squared-error (MSE) law for the decoupled estimator, $$\\textrm{MSE} = (\\tau ^2/m)\\, w^\\top C w + O(m^{-2})$$ , which shows that the reward correlation matrix C sets the achievable gradient- noise floor at fixed score variance and group size, so positively correlated objectives incur a higher noise floor, though the law bounds noise and not signal; and (ii) a closed-form, sign-changing law for the cross-covariance shift induced by gating one reward on another (“ b counts only if a passes”), derived for raw rewards, with an explicit threshold $$\\gamma ^\\star $$ whose sign locates the shift but does not by itself establish an optimization benefit, together with a zero-fill implementation that preserves the U-statistic structure so the MSE law carries over. We additionally give a lattice resolution bound that formalizes GDPO’s “reward signal collapse,” and restate MO-GRPO’s influence law as background: decoupling makes channel influence equal to $$(Cw)_\\ell $$ , which is free of the channel standard deviations for every C and weight-proportional precisely when $$Cw = w$$ . Every claim is verified in three tiers: a synthetic harness that matches each closed form to three significant figures (and falsified two first-draft conjectures), real GSM8K rollouts on Qwen2.5-1.5B and -7B (realized-to-predicted ratios in [0.83, 1.03], with seven of eight $$95\\%$$ bootstrap confidence intervals covering 1.0), and a multi-reward GRPO fine-tuning study on a synthetic fintech domain. The measured effect of decoupling is a reallocation of influence across reward channels, together with the gradient-noise behaviour above. It is not an improvement in aggregate reward: the aggregate-reward difference between the decoupled and scalarized arms is not statistically significant, whereas both per-channel allocation differences are significant at the one percent level. All datasets, fine-tuned models, and code are released, with the construction, reward definitions, partitioning and provenance of each released dataset set out in full in the section on dataset construction.
Authors
- Yiqiao Yin (ORCID: https://orcid.org/0000-0003-1216-4232)
Institutions
- Columbia University (US)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1038/s41598-026-70810-0
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- article
- Field-Weighted Citation Impact
- 0.00