On-Policy Residual Learning for Diffusion Language Models
Abstract
Residual context is emerging as a key ingredient for unlocking the speed–quality Pareto frontier of diffusion language models (DLMs) by carrying forward predictive distributions across denoising steps. Yet turning this seemingly simple mechanism of computation recirculation into an effective and efficient learning paradigm is far from simple: there is no clear recipe for how residual context should be learned. We introduce ORL, a framework for on-policy residual learning in DLMs. Algorithmically, we identify that residual learning is inherently on-policy. ORL learns directly from these model-induced residual trajectories on the fly and introduces an entropy-contraction objective that encourages sharpening predictions as the model revisits its own intermediate predictions. System-wise, exact residual supervision requires capturing, retaining, and repeatedly processing many trajectory-dependent intermediate states. ORL co-designs the learning algorithm with asynchronous residual capture and transport, specialized block-diffusion attention, and packed compute-balanced training, making on-policy residual learning practical at scale. On a 4B DLM converted from Qwen3-4B, ORL improves tokens per forward (TPF) by up to 6.2× and end-to-end throughput by approximately 5× at comparable or better generation quality. On a single H100, ORL-trained models achieve a stronger speed–quality Pareto frontier than autoregressive models accelerated with EAGLE3 and DFlashV1 speculative decoding, under both strict and relaxed verification. On LLaDA-2.1-mini, 36 H100-hours and 4.2M self-generated tokens, about 10× less post-training data than prior LLaDA parallel-decoding recipes, yield up to 1.5× higher TPF and up to 3× higher throughput at matched quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.