Step-Level Gradient Masking for GRPO: Selective Optimization of Reasoning Trajectories

Group Relative Policy Optimization (GRPO) has recently emerged as an effective algorithm for Reinforcement Learning from Verifiable Rewards (RLVR), allowing for improvements in logical reasoning, more specifically in mathematical reasoning in large language models (LLMs) without supervised reasoning traces. It does, however, face a credit assignment problem, in that it applies gradient updates uniformly across every token, including trivial formatting steps and already-mastered reasoning prefixes, thereby wasting gradient capacity. We propose a step-level gradient masking framework for GRPO that restricts policy updates to structured reasoning steps where the model is still actively uncertain, instead of updating all tokens uniformly. The framework decomposes masking into three components: a step-generation protocol, a masking function, and a routing policy driven by a calibrated threshold. We instantiate the framework with two routing signals: KL-GRPO (KL-Divergence GRPO), which routes via per-step KL divergence from a reference model, and EA-GRPO (Entropy-Aware GRPO), which routes via the policy's own Shannon entropy and requires no reference model pass. Training Qwen3-4B on 2,000 stratified RLVR problems with Low-Rank Adaptation (LoRA), EA-GRPO achieves 72.94% on MATH-500 (+1.04% over standard GRPO, paired bootstrap p = 0.0423) and 15.83% on AIME 2025 (+2.83%), while KL-GRPO attains the smoothest gradient trajectory and lowest format error rate (3.4%). These results establish step-level gradient masking as a principled, low-overhead design point for GRPO-family training that improves generalization on hard mathematical reasoning tasks.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23169159
Primary Topic
Reinforcement Learning in Robotics
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Step-Level Gradient Masking for GRPO: Selective Optimization of Reasoning Trajectories

Himani S. Deshpande, Mahesh Patel, Youhan Lalwani
Zenodo (CERN European Organization for Nuclear Research)
Reinforcement Learning in Robotics
preprint

Step-Level Gradient Masking for GRPO: Selective Optimization of Reasoning Trajectories

Himani S. Deshpande, Mahesh Patel, Youhan Lalwani
preprint en

Abstract

Group Relative Policy Optimization (GRPO) has recently emerged as an effective algorithm for Reinforcement Learning from Verifiable Rewards (RLVR), allowing for improvements in logical reasoning, more specifically in mathematical reasoning in large language models (LLMs) without supervised reasoning traces. It does, however, face a credit assignment problem, in that it applies gradient updates uniformly across every token, including trivial formatting steps and already-mastered reasoning prefixes, thereby wasting gradient capacity. We propose a step-level gradient masking framework for GRPO that restricts policy updates to structured reasoning steps where the model is still actively uncertain, instead of updating all tokens uniformly. The framework decomposes masking into three components: a step-generation protocol, a masking function, and a routing policy driven by a calibrated threshold. We instantiate the framework with two routing signals: KL-GRPO (KL-Divergence GRPO), which routes via per-step KL divergence from a reference model, and EA-GRPO (Entropy-Aware GRPO), which routes via the policy's own Shannon entropy and requires no reference model pass. Training Qwen3-4B on 2,000 stratified RLVR problems with Low-Rank Adaptation (LoRA), EA-GRPO achieves 72.94% on MATH-500 (+1.04% over standard GRPO, paired bootstrap p = 0.0423) and 15.83% on AIME 2025 (+2.83%), while KL-GRPO attains the smoothest gradient trajectory and lowest format error rate (3.4%). These results establish step-level gradient masking as a principled, low-overhead design point for GRPO-family training that improves generalization on hard mathematical reasoning tasks.

Zenodo (CERN European Organization for Nuclear Research)
University of Mumbai (IN)
Reinforcement Learning in Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Step-Level Gradient Masking for GRPO: Selective Optimization of Reasoning Trajectories — Himani S. Deshpande, Mahesh Patel, et al. · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS