Stability-Weighted Decoding for Diffusion Language Models
Abstract
Diffusion large language models (dLLMs) enable parallel text generation by iteratively denoising a fully masked sequence. However, existing decoding strategies rely on single-step static confidence metrics, ignoring temporal history and often prematurely unmasking unstable tokens. Under idealized ground-truth context revelation, we establish that the expected KL divergence between true conditional distributions equals the information gained about a token and lower-bounds its dependence on the previously masked context. Inspired by this result, we propose Stability-Weighted Decoding (SWD), a training-free, plug-and-play strategy that incorporates temporal stability into token scoring, using model-based temporal KL as an empirical sensitivity signal rather than a dependency guarantee. Experiments on code generation and mathematical reasoning benchmarks demonstrate that SWD improves generation accuracy across representative scoring metrics and exhibits exceptional robustness under aggressive acceleration ratios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.