Guiding Discrete Diffusion with Distilled Process Rewards
Abstract
Masked diffusion language models such as LLaDA generate text by iteratively committing tokens from a fully masked canvas. An erroneous early commitment can influence subsequent denoising steps and lead to an incorrect final answer. Process reward models (PRMs) can localize these errors, but running a separate 7B-parameter PRM at inference is expensive. We present a method for distilling process rewards into a frozen diffusion language model. A value head summarizes the quality of a partially masked state and supports post-hoc repair; a step-issue head predicts token-level issue scores and guides commitment order during the original denoising loop. We evaluate a single 256-token canvas to isolate adaptive commitment ordering from fixed block boundaries: lower-risk positions can be committed in parallel, while higher-risk positions wait for additional context. In one reported GSM8K run with LLaDA-8B-Instruct, step-issue guidance changes accuracy from 44.7% to 53.9% (+9.2 pp) with a 1.1M-parameter MLP. Matched controls in that run do not reproduce this change with random or shuffled issue scores. On MATH-500, the corresponding GSM8K-to-MATH-500 transfer run changes accuracy from 24.0% to 25.8%. Direct PRM repair reaches a higher point estimate but requires the full PRM at inference and an additional regeneration pass. In a matched 100-example GSM8K timing run, guidance was only 0.4% slower than baseline, versus 23.6% for distilled repair and 35.1% for direct PRM repair.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.