DUET: Dual-Sparse Inference for Diffusion Language Models
Abstract
Diffusion language models (DLMs) enable parallel token generation through iterative denoising, yet their inference remains costly due to repeated attention over a growing historical prefix and repeated processing of the entire current generation block. Historical attention becomes increasingly expensive as generation progresses, while all current-block positions are propagated through the Transformer at every denoising step even though only a small fraction are typically resolved at that step. In this paper, we introduce Duet, a training-free dual-sparse inference framework that addresses these challenges through prefix Key-Value (KV) sparsity and early query sparsity, respectively. Specifically, Duet exploits the concentration and cross-step stability of historical attention within each generation block to construct a compact historical KV subset that is reused across denoising steps. For query sparsity, Duet predicts update positions from shallow representations and propagates only the corresponding queries through the suffix Transformer layers, avoiding deep computation at positions unlikely to be updated in the current step. Experiments show that Duet maintains competitive accuracy while achieving 1.34/2.83 end-to-end/per-step speedups on LLaDA2.1-mini and 2.41/4.72 on SDAR-8B-Chat-b32 at a 32K context length. The anonymized code is available at https://anonymous.4open.science/r/Duet-45AF.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.