acceptodds
Under review as a conference paper at ICLR 2027

RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers

Abstract

Diffusion Transformers (DiTs) have become dominant for high-fidelity video generation, yet their attention cost bottlenecks long-sequence synthesis. Recent sparse-linear hybrids reduce this cost but degrade at extreme sparsity due to what we call the RoPE dilemma: positive feature maps enabling linear attention do not preserve 3D Rotary Position Embeddings (RoPE) orthogonality, losing explicit relative spatiotemporal offsets in global aggregation. We propose RoPeSLR, a 3D RoPE-aware sparse-low-rank framework. Through structural analysis, we observe that DiT attention matrices empirically exhibit decomposition into high-energy sparse spikes and smooth backgrounds with slowly growing rank. Motivated by this observation, we design an energy-driven block-sparse branch for dominant interactions and a Fourier low-rank compensator preserving RoPE structure via rotated Fourier features aggregating global context in . At 90% sparsity, RoPeSLR maintains strong VBench quality on Wan2.1-1.3B and 14B while reducing DiT denoising latency by and , respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.