DepCap: Adaptive Block-Wise Parallel Decoding for Efficient Diffusion LM Inference
Abstract
Diffusion language models (DLMs) offer a promising alternative to autoregressive language generation by enabling parallel decoding and global sequence refinement. To unlock this potential, DLM inference must balance generation quality and decoding speed. Recent block-wise DLM decoding methods improve this trade-off by performing diffusion-based decoding sequentially in blocks. However, existing block-wise methods often rely on signals that are not well aligned with two key decisions in DLM inference. For block partitioning, fixed schedules and current-step local signals do not explicitly capture cross-step prediction changes, while confidence-based parallel decoding does not explicitly consider interactions between token predictions. In this paper, we argue that block-wise DLM inference requires decision-matched signals: cross-step signals for block partitioning and token-level conflict signals for parallel decoding. Based on this view, we propose DepCap, a training-free framework that instantiates the cross-step signal as the influence of the last decoded block and uses it to adaptively determine the next block boundary. Within each block, DepCap combines token confidence with a pairwise conflict score to select tokens for parallel decoding, balancing decoding speed and generation quality. DepCap is plug-and-play across DLMs and compatible with existing block-wise KV-cache strategies. An information-theoretic analysis characterizes the deviation from additive last-block influence, motivating the proposed block-partitioning criterion. Experiments across four DLM backbones and four reasoning and coding benchmarks show favorable quality-speed trade-offs, with up to 5.63 speedup over vanilla decoding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.