Every Step Counts in RL for Diffusion Language Models
Abstract
Diffusion language models are a promising and more efficient alternative to au- toregressive models. However, most RL post-training techniques are not easily adapted to these models. One issue is that diffusion language models generate answers through many intermediate states, deciding at each step which masked positions to reveal and which tokens to place in them. Many existing RL methods used in this setting approximate intractable sequence likelihoods or weight every denoising step by the same advantage, without explicitly distinguishing how these intermediate states of the denoising trajectory affect the expected reward. To ad- dress this issue of trajectory blindness, we introduce PACT (Potential-Anchored Consistency of Transitions). Motivated by a reward-tilted reference distribution, PACT jointly learns the policy and a scalar potential that scores each intermedi- ate state, anchored by terminal reward. For each sampled denoising transition, a consistency loss matches the policy-to-reference log-probability ratio of revealed tokens to the change in this potential. This formulation requires neither sequence- likelihood estimation nor importance-ratio clipping. On four tasks, PACT improves LLaDA-8B-Instruct by +9.9 points on GSM8K, +12.0 on MATH500, +55.7 on Countdown, and +72.6 on Sudoku, and reaches the best average accuracy across the four tasks reported for this model, +3.7 points over the next best method. Further compared to AGRPO, it reaches that baseline’s best accuracy in about half the training compute on GSM8K.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.