HorizonKV: An Overlap-Aware Sparse Attention System for Long-Horizon Autoregressive Video Diffusion
Abstract
Autoregressive diffusion transformers (AR-DiTs) generate video chunk by chunk and can in principle extend generation to arbitrary video length. However, as KV caches grow, existing models are forced to retain only a short window of local frames, causing catastrophic forgetting and drift. Retrieval-based sparse attention retrieves important historical frames, but GPU capacity limits the range of accessible history. We observe that access to the full KV history substantially improves generation quality. However, realizing this benefit raises two challenges. The first is insufficient memory capacity. Recent work addresses this by offloading KV to the CPU, but retrieval overhead slows generation. The second is increased compute FLOPs, as longer histories contain more important KV entries and thus require more attention FLOPs. To address these challenges, we present HorizonKV, a training-free sparse attention system for long-horizon autoregressive video diffusion. To address insufficient memory capacity, we introduce Overlap-Aware Execution, which uses tiered storage and selectively overlaps KV transfers with attention computation to hide transfer latency. To address increased compute FLOPs, we introduce Residency-Aware Sparse Attention, which adaptively selects two-level sparsity configurations based on KV cache residency on the GPU. It maintains quality while hiding as much communication overhead as possible without increasing FLOPs. Experiments show HorizonKV advances the Pareto frontier over all baselines, with up to 1.486x end-to-end speedups.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.