Value Gradient Hypothesis of RL for LLMs
Abstract
Reinforcement learning substantially improves pretrained language models, but it remains understudied why critic-free methods such as GRPO and critic-free PPO variants work as well as they do, and when they should provide the largest gains. We develop a value-gradient perspective of critic-free RL for LLM post-training. First, under a differentiable rollout and additive-noise parameterization, we identify when the expected actor update admits a pathwise representation: the corresponding return costates have conditional expectation equal to the value gradient. Second, for discrete transformer policies, we characterize empirical adjoints through the layer-position computation graph and bound their discrepancy from a specified relaxed value signal by local-signal and transition residuals. Policy entropy controls one relaxation component of this bound, not the entire discrepancy. These results motivate a hypothesis decomposing RL impact into value-gradient signal and reachable reward headroom, yielding a criterion for when RL may be most effective along a pretraining trajectory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.