acceptodds
Under review as a conference paper at ICLR 2027

Value Gradient Hypothesis of RL for LLMs

Abstract

Reinforcement learning substantially improves pretrained language models, but it remains understudied why critic-free methods such as GRPO and critic-free PPO variants work as well as they do, and when they should provide the largest gains. We develop a value-gradient perspective of critic-free RL for LLM post-training. First, under a differentiable rollout and additive-noise parameterization, we identify when the expected actor update admits a pathwise representation: the corresponding return costates have conditional expectation equal to the value gradient. Second, for discrete transformer policies, we characterize empirical adjoints through the layer-position computation graph and bound their discrepancy from a specified relaxed value signal by local-signal and transition residuals. Policy entropy controls one relaxation component of this bound, not the entire discrepancy. These results motivate a hypothesis decomposing RL impact into value-gradient signal and reachable reward headroom, yielding a criterion for when RL may be most effective along a pretraining trajectory.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.