Rethinking Sparse Attention in Few-Step Diffusion Transformer
Abstract
Few-step diffusion compresses a multi-step teacher trajectory into a handful of student evaluations. Sparse attention further reduces cost by retaining only highly ranked interactions, but the student’s current ranking need not represent the corresponding teacher-window support. On aligned probes of a four-step Wan2.2 student, the top-b pool captures 51.91% of teacher attention mass and leaves a value-weighted residual. The useful core favors persistent interactions; its complement contains both missed salient blocks and a diffuse tail. This diagnosis motivates SPARE Attention: current-score candidates plus a small complement quota under a fixed sparse budget, with coarse routing and no teacher access at inference. A matched per-step-budget intervention reduces final Teacher40 RMSE from 0.959537 to 0.928340 (3.25%), improving all six reported records; larger quotas are not uniformly better. Separately, the original system configuration selects 14.84% of key blocks on the full 6,220-video VBench suite and denoises in 29.565 s, the shortest measured time among the timed sparse implementations, with Total comparable to the SLA family and PISA
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.