acceptodds
Under review as a conference paper at ICLR 2027

Decouple to Decode: Parallel Decoding with Deferred Proposal Verification for Diffusion Language Models

Abstract

Discrete diffusion language models (dLLMs) alleviate the sequential bottleneck of autoregressive generation by predicting multiple masked tokens in parallel at each denoising step. However, existing parallel decoding methods typically make single-step commitment decisions, coupling whether a prediction is reliable enough to commit with whether it is worth retaining for further evidence, thereby tying broader proposal coverage to looser commitment criteria. Indeed, we observe two complementary regimes: one where predictions highly agree with the final tokens and are suitable for immediate commitment, and another where candidates of varying reliability remain mixed but become more separable under cross-step re-evaluation. Motivated by this observation, we propose Decouple to Decode (D2D), a training-free framework that decouples immediate commitment from deferred verification. D2D uses conditional risk budgets to commit a Core and retain a masked Shell for one-step verification; in the next ordinary forward pass, Shell proposals are accepted if their token identities remain unchanged and confidence does not decrease, requiring no extra verifier or model evaluation. Experiments on LLaDA and Dream across five benchmarks spanning mathematical reasoning, code generation, and instruction following demonstrate up to wall-clock speedup over standard decoding and % over confidence decoding, while maintaining comparable task performance. Combined with KV caching, D2D achieves up to speedup over standard decoding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.