Supporting Tokens First: Dirichlet Prior-Guided Decoding for Diffusion Language Models
Abstract
Unlike auto-regressive language models, diffusion language models (DLMs) iteratively refine masked sequences and can reveal multiple positions in parallel, offering the potential for faster generation. However, unresolved positions can remain conditionally dependent, making parallel decoding potentially unreliable. Explicitly modelling these dependencies through latent-variable inference is computationally expensive, while existing DLM decoding strategies primarily rely on instance-level signals, such as model confidence or attention, which can be noisy at capturing a token's influence on subsequent predictions. To address these limitations, we hypothesise the existence of supporting tokens that, when revealed early, reduce conditional dependence among the remaining masked positions. Under an idealised Bayesian latent-variable model, we show that prioritising tokens with greater prior support can reduce residual dependence in expectation. Building on this insight, we propose Prior-Guided Diffusion Decoding (PGDD), a training-free framework that exploits stable, population-level vocabulary statistics to identify candidate supporting tokens without costly latent-variable inference. Specifically, an answer-free Dirichlet Prior Bank captures query-conditioned vocabulary priors from an external corpus, which are combined with the DLM's current predictions through a Dirichlet posterior update to guide token commitment. By complementing noisy instance-level confidence with population-level prior information, PGDD prioritises supporting tokens, reducing residual conditional dependence and enabling more reliable parallel decoding. Empirical analysis confirms that higher-prior tokens tend to yield greater confidence gains at other masked positions early in denoising, while prior-guided ordering reduces residual conditional dependence. Experiments on mathematical reasoning, code generation, and planning demonstrate improved generation quality and decoding efficiency with limited computational overhead compared with recent baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.