SparkDiffusion: A Unified Sparse Warm-Up and Few-Step Distillation Framework for Mitigating the High-Sparsity Trap
Abstract
Video diffusion transformers are central to visual generation, but inference is dominated by attention over long spatiotemporal sequences. While video attention's sparse-plus-low-rank structure suggests high sparsity is possible, existing models degrade beyond sparsity even after step-local training converges. We identify this failure regime as the high-sparsity trap: step-local objectives leave terminal-visible velocity errors that accumulate during sampling without direct gradient signals to correct them. We propose \method (\methodfull), a unified framework whose central principle is to first adapt a dense video DiT into a high-sparsity coarse prior through sparse warm-up, and then correct its terminal trajectory through few-step trajectory-mixed distillation; fused FP8 kernels provide the deployment layer. Our idealized analysis isolates the mechanism and limits, while experiments validate the staged framework across multiple backbones, tasks, resolutions, and hardware platforms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.