acceptodds
Under review as a conference paper at ICLR 2027

AsyncLane: Decoupling Refinement from Advancement in Diffusion Language Model Decoding

Abstract

Block decoding in diffusion language models typically starts the next block only after the current block is fully completed, letting a few pending positions delay subsequent advancement. In the serial decoding trajectories we analyze, some next-block predictions remain consistent before and after current-block completion. Our key insight is that completion must resolve all pending positions whereas advancement can begin from a subset of candidates, so the two can be decoupled. Based on this insight, we propose AsyncLane, a training-free scheduling method for block decoding. The main lane continues completing the current block while the private lane accumulates decoding progress for the next block. During overlap, both are computed in a single batched forward pass, and private progress is handed off after the current block completes. In LLaDA and Dream, private progress is retained as candidate tokens that support subsequent predictions in the private view; lookahead prefill (LA) further reuses full-sequence predictions produced at opening. We evaluate AsyncLane on LLaDA, Dream, and DiffusionGemma using five benchmarks spanning mathematical reasoning, code generation, and instruction following. Under the reported configurations, effective output throughput on LLaDA and Dream reaches up to 2.35 that of AdaBlock, and on DiffusionGemma up to 1.23 that of the official decoder. Quality changes are task-dependent; on multiple tasks quality and throughput improve jointly—for example, Dream's accuracy on MATH500 is 3.80 percentage points higher than AdaBlock's. Controlled ablations on LLaDA and Dream further support practical request-latency gains from early opening and show that full LA provides an additional latency reduction under fixed early opening.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.