A Method-of-Types Interpretation of Token Distribution Convergence for Denoising Step Reduction in Uniform-State Diffusion LLMs
Abstract
Diffusion LLMs (dLLMs) alleviate the memory bottleneck of autoregressive LLMs through block-wise parallel generation, but their generation speed remains limited by the number of denoising steps. Existing samplers either ignore convergence across consecutive steps or employ conservative criteria, leaving already-converged positions unexploited and causing redundant steps. We introduce the **Convergence Probability**, which estimates the likelihood that consecutive-step probability distributions have converged through a probabilistic interpretation based on the method of types in Shannon information theory. Based on this quantity, we propose the **Convergence-Aware Entropy-Bounded Sampler**, a training-free sampler for uniform-state dLLMs that prioritizes converged positions while considering both confidence and joint dependency. Across reasoning, coding, and instruction-following benchmarks, existing methods range from a 78.2% increase to a 7.0% reduction in denoising steps, whereas our method consistently reduces them by 9.9%–28.9%. Further, our experiments show that accepting converged positions accelerates denoising by reducing the entropy of neighboring positions, thereby rapidly resolving uncertainty across the entire canvas.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.