HiAda: History-Aware Adaptive Decoding for Diffusion Large Language Models
Abstract
Diffusion large language models (dLLMs) enable parallel token generation through iterative denoising, but their practical inference remains costly due to dynamically changing bidirectional representations, the limited applicability of conventional KV caching, and rigid block-wise decoding schedules. These limitations call for inference strategies that can reuse computation while adapting to the evolving decoding state. We propose HiAda, a training-free, history-aware, and adaptive inference framework that exploits predictions and attention signals accumulated throughout the denoising trajectory. HiAda constructs a token-level active set to jointly coordinate KV-cache updates and token commitment, while dynamically adapting the decoding scope to the model's evolving confidence. This design reduces redundant computation, mitigates stale cache reuse, and improves the effective parallelism of dLLM decoding. Experiments on LLaDA-8B-Instruct and Dream-7B-Base across GSM8K, MATH, HumanEval, and MBPP show that HiAda achieves up to 51.1× speedup over vanilla diffusion decoding and up to 2.01× speedup over Fast-dLLM, while offering a controllable quality–efficiency trade-off. These results demonstrate that combining decoding history with adaptive computation can substantially improve the efficiency of diffusion language model inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.