Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
Abstract
Diffusion Large Language Models (dLLMs) offer a compelling paradigm for natural language generation, leveraging parallel decoding and bidirectional attention to achieve superior global coherence compared to autoregressive models. While recent works have accelerated inference via KV cache reuse or heuristic decoding, they overlook the intrinsic inefficiencies within the block-wise diffusion process. Specifically, these methods suffer from spatial redundancy by modeling sparse suffix regions uniformly and temporal inefficiency by applying fixed denoising schedules across decoding. To address this, we propose Streaming-dLLM, a training-free framework that streamlines inference across both spatial and temporal dimensions. Spatially, we introduce approximated suffix pruning that leverages context modeling redundancy to compress suffix cost from linear to constant complexity, decoupling inference cost from sequence length. Temporally, we employ a dynamic confidence aware strategy that adaptively modulates denoising thresholds based on global mask state evolution, allowing the model to skip unnecessary iterations for converged tokens. Experiments on five text-only dLLMs and one diffusion multimodal large language model (dMLLM) across six benchmarks show that Streaming-dLLM matches or improves the accuracy of existing acceleration methods while achieving speedups of up to over these methods. For long sequence generation, it outperforms the strongest baseline by 2.5 accuracy points with higher throughput.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.