Step-Level Gradient Masking for GRPO: Selective Optimization of Reasoning Trajectories
Group Relative Policy Optimization (GRPO) has recently emerged as an effective algorithm for Reinforcement Learning from Verifiable Rewards (RLVR), allowing for improvements in logical reasoning, more specifically in mathematical reasoning in large language models (LLMs) without supervised reasoning traces. It does, however, face a credit assignment problem, in that it applies gradient updates uniformly across every token, including trivial formatting steps and already-mastered reasoning prefixes, thereby wasting gradient capacity. We propose a step-level gradient masking framework for GRPO that restricts policy updates to structured reasoning steps where the model is still actively uncertain, instead of updating all tokens uniformly. The framework decomposes masking into three components: a step-generation protocol, a masking function, and a routing policy driven by a calibrated threshold. We instantiate the framework with two routing signals: KL-GRPO (KL-Divergence GRPO), which routes via per-step KL divergence from a reference model, and EA-GRPO (Entropy-Aware GRPO), which routes via the policy's own Shannon entropy and requires no reference model pass. Training Qwen3-4B on 2,000 stratified RLVR problems with Low-Rank Adaptation (LoRA), EA-GRPO achieves 72.94% on MATH-500 (+1.04% over standard GRPO, paired bootstrap p = 0.0423) and 15.83% on AIME 2025 (+2.83%), while KL-GRPO attains the smoothest gradient trajectory and lowest format error rate (3.4%). These results establish step-level gradient masking as a principled, low-overhead design point for GRPO-family training that improves generalization on hard mathematical reasoning tasks.
Authors
- Himani S. Deshpande (ORCID: https://orcid.org/0000-0002-9050-0732)
- Mahesh Patel
- Youhan Lalwani (ORCID: https://orcid.org/0009-0006-3646-1011)
Institutions
- University of Mumbai (IN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23169160
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- preprint