Decoupling and conditioning reshape influence allocation and the gradient-noise floor in multi-reward GRPO under a finite-sample U-statistic analysis

Abstract Multi-reward reinforcement learning with verifiable rewards (RLVR) increasingly relies on Group Relative Policy Optimization (GRPO). Two recent methods—MO-GRPO and GDPO—replace GRPO’s standard aggregate-then-normalize (AN) advantage with a decoupled normalize-then-aggregate (NA) estimator, and show empirically that AN lets the highest-variance reward channel dominate and collapses distinct reward combinations into identical advantages. We supply the missing finite-sample, correlation-aware theory of these estimators. Our two main results are: (i) an exact finite-sample mean-squared-error (MSE) law for the decoupled estimator, $$\\textrm{MSE} = (\\tau ^2/m)\\, w^\\top C w + O(m^{-2})$$ , which shows that the reward correlation matrix C sets the achievable gradient- noise floor at fixed score variance and group size, so positively correlated objectives incur a higher noise floor, though the law bounds noise and not signal; and (ii) a closed-form, sign-changing law for the cross-covariance shift induced by gating one reward on another (“ b counts only if a passes”), derived for raw rewards, with an explicit threshold $$\\gamma ^\\star $$ whose sign locates the shift but does not by itself establish an optimization benefit, together with a zero-fill implementation that preserves the U-statistic structure so the MSE law carries over. We additionally give a lattice resolution bound that formalizes GDPO’s “reward signal collapse,” and restate MO-GRPO’s influence law as background: decoupling makes channel influence equal to $$(Cw)_\\ell $$ , which is free of the channel standard deviations for every C and weight-proportional precisely when $$Cw = w$$ . Every claim is verified in three tiers: a synthetic harness that matches each closed form to three significant figures (and falsified two first-draft conjectures), real GSM8K rollouts on Qwen2.5-1.5B and -7B (realized-to-predicted ratios in [0.83, 1.03], with seven of eight $$95\\%$$ bootstrap confidence intervals covering 1.0), and a multi-reward GRPO fine-tuning study on a synthetic fintech domain. The measured effect of decoupling is a reallocation of influence across reward channels, together with the gradient-noise behaviour above. It is not an improvement in aggregate reward: the aggregate-reward difference between the decoupled and scalarized arms is not statistically significant, whereas both per-channel allocation differences are significant at the one percent level. All datasets, fine-tuned models, and code are released, with the construction, reward definitions, partitioning and provenance of each released dataset set out in full in the section on dataset construction.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-21
DOI
https://doi.org/10.1038/s41598-026-70810-0
Primary Topic
Reinforcement Learning in Robotics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Decoupling and conditioning reshape influence allocation and the gradient-noise floor in multi-reward GRPO under a finite-sample U-statistic analysis

Yiqiao Yin
Scientific Reports
Reinforcement Learning in Robotics
article

Decoupling and conditioning reshape influence allocation and the gradient-noise floor in multi-reward GRPO under a finite-sample U-statistic analysis

Yiqiao Yin
article en

Abstract

Abstract Multi-reward reinforcement learning with verifiable rewards (RLVR) increasingly relies on Group Relative Policy Optimization (GRPO). Two recent methods—MO-GRPO and GDPO—replace GRPO’s standard aggregate-then-normalize (AN) advantage with a decoupled normalize-then-aggregate (NA) estimator, and show empirically that AN lets the highest-variance reward channel dominate and collapses distinct reward combinations into identical advantages. We supply the missing finite-sample, correlation-aware theory of these estimators. Our two main results are: (i) an exact finite-sample mean-squared-error (MSE) law for the decoupled estimator, $$\textrm{MSE} = (\tau ^2/m)\, w^\top C w + O(m^{-2})$$ , which shows that the reward correlation matrix C sets the achievable gradient- noise floor at fixed score variance and group size, so positively correlated objectives incur a higher noise floor, though the law bounds noise and not signal; and (ii) a closed-form, sign-changing law for the cross-covariance shift induced by gating one reward on another (“ b counts only if a passes”), derived for raw rewards, with an explicit threshold $$\gamma ^\star $$ whose sign locates the shift but does not by itself establish an optimization benefit, together with a zero-fill implementation that preserves the U-statistic structure so the MSE law carries over. We additionally give a lattice resolution bound that formalizes GDPO’s “reward signal collapse,” and restate MO-GRPO’s influence law as background: decoupling makes channel influence equal to $$(Cw)_\ell $$ , which is free of the channel standard deviations for every C and weight-proportional precisely when $$Cw = w$$ . Every claim is verified in three tiers: a synthetic harness that matches each closed form to three significant figures (and falsified two first-draft conjectures), real GSM8K rollouts on Qwen2.5-1.5B and -7B (realized-to-predicted ratios in [0.83, 1.03], with seven of eight $$95\%$$ bootstrap confidence intervals covering 1.0), and a multi-reward GRPO fine-tuning study on a synthetic fintech domain. The measured effect of decoupling is a reallocation of influence across reward channels, together with the gradient-noise behaviour above. It is not an improvement in aggregate reward: the aggregate-reward difference between the decoupled and scalarized arms is not statistically significant, whereas both per-channel allocation differences are significant at the one percent level. All datasets, fine-tuned models, and code are released, with the construction, reward definitions, partitioning and provenance of each released dataset set out in full in the section on dataset construction.

Scientific Reports
Columbia University (US)
Decent work and economic growth
Openalex Percentile: Top 8%
Reinforcement Learning in Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.