Loop-Distill: Refining Information Flow in Diffusion Large Language Models
Abstract
Diffusion language models (DLMs) reduce decoding latency by denoising multiple token positions in parallel. A practical way to obtain this is to convert a pretrained autoregressive (AR) large language model into a DLM with the block diffusion paradigm. However, this conversion also changes the role of a forward pass. Instead of predicting one next token from a fully observed prefix, the model must form and refine representations for multiple masked positions under partial context. We show that this mismatch gives rise to an information flow bottleneck: masked token states stabilize prematurely in the shallow layers, yet the predictions made by these layers are often overconfident and rarely consistent with the final output. To address this, we propose _Loop-Distill_, which uses local shallow-layer loop refinement to let masked states absorb known context more fully and distills this deeper shallow refinement into the original single-pass model. During training, we exploit DLM decoding as an intrinsic refinement loop by constructing an online stop-gradient teacher from the student’s current weights and looping selected shallow layers to refine masked states. This compresses useful loop computation into the deployed model without changing its inference-time architecture. Across 13 benchmarks, Loop-Distill improves average accuracy on GLM-4.7-Flash and Qwen-30B-A3B-2507 by 2.05 and 2.21 points, respectively. It also improves decoding throughput by 11.0% on average by enabling more effective masked state refinement and thus fewer full-model denoising passes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.