acceptodds
Under review as a conference paper at ICLR 2027

Measuring Gradient Geometry in Language-Model RLVR with Sketched Gradients

Abstract

Reinforcement learning with verifiable rewards (RLVR) has emerged as a dominant paradigm in training large language models to solve a wide range of tasks, including mathematical reasoning. Most RLVR pipelines build on a variant of the standard REINFORCE policy-gradient algorithm. Many methods make claims about the noise of policy-gradient estimators, and about the relative geometry between policy gradients on different prompts: for example, baselines as control variates, and pass@ objectives to shift gradient weight from easier to harder problems. However, to our knowledge, no prior work has directly measured these quantities in language-model RLVR, at scale and across training. We compress per-sequence gradients of online rollouts with CountSketch, a linear operation which allows offline analysis to measure the effect of RLVR objectives, rollout allocations and baselines on policy-gradient estimation. After validating that sketches preserve gradient geometry better than common proxies, such as the gradient through the language-model head, we use them to study within-prompt noise and cross-prompt structure. Per-prompt policy-gradient estimates are noisy: at 16 rollouts per prompt, the median cosine similarity between an estimate and a held-out, high-budget estimate for the same prompt is between 0.02 and 0.20, depending on model and checkpoint. We find that, at the beginning of training, prompts share a common policy-gradient direction; at the reward plateau, per-prompt gradients are easier to estimate, but full-batch gradients become more difficult to estimate as the early shared cross-prompt direction disappears. Further, we show that group-relative baselines typically produce more reliable policy-gradient estimates than fixed baselines, that allocating rollouts by estimated pass rate improves reliability over uniform allocation, and that pass@ gradients that shift weight toward harder prompts are harder to estimate as grows. We propose this sketched gradient analysis as a tool to measure the effect of various algorithmic choices in RLVR.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.