acceptodds
Under review as a conference paper at ICLR 2027

Loop-Distill: Refining Information Flow in Diffusion Large Language Models

Abstract

Diffusion language models (DLMs) reduce decoding latency by denoising multiple token positions in parallel. A practical way to obtain this is to convert a pretrained autoregressive (AR) large language model into a DLM with the block diffusion paradigm. However, this conversion also changes the role of a forward pass. Instead of predicting one next token from a fully observed prefix, the model must form and refine representations for multiple masked positions under partial context. We show that this mismatch gives rise to an information flow bottleneck: masked token states stabilize prematurely in the shallow layers, yet the predictions made by these layers are often overconfident and rarely consistent with the final output. To address this, we propose _Loop-Distill_, which uses local shallow-layer loop refinement to let masked states absorb known context more fully and distills this deeper shallow refinement into the original single-pass model. During training, we exploit DLM decoding as an intrinsic refinement loop by constructing an online stop-gradient teacher from the student’s current weights and looping selected shallow layers to refine masked states. This compresses useful loop computation into the deployed model without changing its inference-time architecture. Across 13 benchmarks, Loop-Distill improves average accuracy on GLM-4.7-Flash and Qwen-30B-A3B-2507 by 2.05 and 2.21 points, respectively. It also improves decoding throughput by 11.0% on average by enabling more effective masked state refinement and thus fewer full-model denoising passes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.