acceptodds
Under review as a conference paper at ICLR 2027

Confidence-Adaptive Permutation-Tree Self Speculative Decoding for Diffusion Large Language Models

Abstract

Diffusion-based Large Language Models (dLLMs) offer a promising alternative to autoregressive models by enabling parallel decoding and bidirectional context modeling. However, existing speculative decoding methods for dLLMs rely on linear or heuristic verification structures that fail to capture the full space of valid generation orders within a block, leading to premature rejection of correct tokens and suboptimal acceptance rates. In this work, we propose CAPSpec (Confidence- Adaptive Perm-tree Speculative Decoding), a training-free framework targeting block-based diffusion models. CAPSpec exploits the key insight that, within a fixed-length block, the n remaining masked positions admit n! possible decoding orders, forming a full permutation tree whose size grows factorially. By first constructing this complete permutation tree over a block’s decoding trajectories and then applying Adaptive Confidence Partitioning (ACP)—which clusters candidate positions by their draft-stage confidence and prunes low-value branches before verification—CAPSpec compresses the factorial-sized search space into a compact tree that retains all high-probability trajectories. A single batched forward pass over the pruned tree evaluates the retained trajectories and identifies a unique accepted path. Extensive experiments on LLaDA-8B-Instruct across mathematics, code, and dialogue benchmarks demonstrate that CAPSpec achieves up to 3.08× wall-clock speedup over step-by-step baselines, establishing a new state of the art for training-free speculative decoding in block-based diffusion language models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.