acceptodds
Under review as a conference paper at ICLR 2027

Enabling Reparameterized Policy Gradients for LLM Post-training

Abstract

The score function (SF) estimator serves as the basis for nearly all reinforcement learning (RL) pipelines currently used to finetune Large Language Models (LLMs). While general, these methods rely exclusively on samples to derive an estimate of the policy gradient, which leads to a limited ability to explore beyond a base policy's distribution. Conversely, one could elect to directly differentiate through a reparameterized objective to fully leverage a reward model's gradient towards improvement rather than estimating it through samples. That said, reparameterized policy gradients have yet to be explored in this area because of the added complexity of differentiating through a discrete sampling objective. In addition, these types of gradients have historically been avoided by RL practitioners, due to their instability and impracticality in long horizon tasks involving auto-regressive generation. In this paper, we challenge these concerns and provide a principled framework for efficiently computing the reparameterized policy gradient for finetuning LLMs with RL. In a series of experiments spanning both preference alignment and reasoning, we show that directly backpropagating through a relaxed objective is not only feasible, but can actually be competitive with current state of the art SF-based estimators despite the literature's warnings against it. Finally, we highlight the unique property of reparameterized policy gradients to explore novel solutions outside the initial sampling distribution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.