Temporal Delta Activation Sparsity for Diffusion Language Model
Abstract
Diffusion language models repeatedly process a sequence to generate text, making inference expensive. Existing activation-sparsity methods reduce this inference cost by discarding small activations. We instead exploit the fact that activations often change little between denoising steps. Thus, we propose Temporal Delta Sparsity (TDS), a training-free method that updates each linear layer using the selected changes to a stored reference input. A shared budget across tokens and channels assigns computation to the changes with the largest weight-informed scores. We show that the accumulator remains consistent with its reference in exact arithmetic and that this selection minimises a one-step output-error bound. Moreover, unprocessed changes remain in the residual for later steps. On GSM8K, MATH500, HumanEval and MBPP, TDS skips – of linear-layer multiply–accumulates. Its average score is versus for dense LLaDA-8B, and versus for dense Dream-7B. On the separate GSM8K evaluation, a whole-row variant achieves and end-to-end speedups, respectively, with equal or higher observed accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.