acceptodds
Under review as a conference paper at ICLR 2027

SparkDiffusion: A Unified Sparse Warm-Up and Few-Step Distillation Framework for Mitigating the High-Sparsity Trap

Abstract

Video diffusion transformers are central to visual generation, but inference is dominated by attention over long spatiotemporal sequences. While video attention's sparse-plus-low-rank structure suggests high sparsity is possible, existing models degrade beyond sparsity even after step-local training converges. We identify this failure regime as the high-sparsity trap: step-local objectives leave terminal-visible velocity errors that accumulate during sampling without direct gradient signals to correct them. We propose \method (\methodfull), a unified framework whose central principle is to first adapt a dense video DiT into a high-sparsity coarse prior through sparse warm-up, and then correct its terminal trajectory through few-step trajectory-mixed distillation; fused FP8 kernels provide the deployment layer. Our idealized analysis isolates the mechanism and limits, while experiments validate the staged framework across multiple backbones, tasks, resolutions, and hardware platforms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.