acceptodds
Under review as a conference paper at ICLR 2027

dSpec: Efficient Speculative Decoding for Block-Diffusion Language Models

Abstract

Speculative decoding accelerates a target language model by letting a lightweight drafter propose several tokens that the target verifies together. For block-diffusion language models, which decode one block of tokens at a time, it offers an alternative to the prevailing approach of lowering a confidence threshold, which trades generation quality for speed. Existing speculative decoders for these models, however, use trajectory-based or autoregressive-style verification, whose draft quality and continuation-state costs both affect acceleration. We present dSpec, which rests on the observation that a target pass over a fully masked block already predicts every position and thus yields a complete draft. dSpec uses this draft directly or corrects it with a one-layer head that reads the target's representation at each position. A single folded target pass then verifies the draft while preparing the block's key–value state and the next block's draft, so a fully accepted block costs only one pass. Requests rejected at different positions are recovered jointly in one shared pass. In SGLang, dSpec achieves 1.6–4.3× the throughput of Fast-dLLM's threshold-1.0 decoding on LLaDA2.0-mini and SDAR-8B at batch sizes 1–128. At batch one, its seven-task mean quality is within about one percentage point of that baseline. On matched hardware, it is roughly twice as fast as the fastest of three existing speculative decoders. Our analysis further shows that longer acceptance does not guarantee higher throughput, so proposal quality and continuation cost must be optimized together.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.