acceptodds
Under review as a conference paper at ICLR 2027

Video DeltaNet: A Video-Native Hybrid Attention for Efficient Video Generation

Abstract

Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for long-range video context. Its linear branch introduces Video Delta Attention (VDA), which jointly resolves correlated within-frame writes and yields a non-expansive inherited-state transition without additional frame-size key scaling. Separate output projections and learnable gates calibrate the two branches, while a staged teacher-alignment recipe progressively introduces the new pathway into pretrained models. We instantiate VDN on MiniMax H3, applying the hybrid to video-to-video interactions while retaining Softmax for interactions involving text or audio. Applying VDN directly yields 2.6× and 3.2× backbone speedups on B200 and H200, respectively. With 8-NFE distillation and optimized inference, VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, corresponding to a 14.5× speedup over the 50-NFE Dense H3 baseline on the same GPU count.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.