HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation
Abstract
Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attention the dominant inference cost. This cost is governed by how many tokens sit in the cache, so it becomes the bottleneck precisely in the regime the field is moving toward: high resolution, where each frame contributes many tokens, and long horizons, where frames accumulate. Existing remedies either evict the cache with coarse heuristics that cause inter-frame flickering, or require model re-training. We show that the attention heads of a pre-trained AR video DiT possess a stable structural property: each head's spatio-temporal receptive field is reproducible from a small, fixed sub-context, and which sub-context suffices can be identified from a single forward pass. We formalize this as four archetypes – Sink, Dummy, Spatial, and Global – and build HeadCast, a training-free, plug-and-play framework that classifies every head once at the maximum-noise step and restructures the monolithic KV cache into head-specific pathways. Crucially, it retains the Global heads that preserve the long-range temporal consistency aggressive eviction destroys. Because the Spatial pathway operates on a fixed-size grid, its savings grow monotonically with the token count: across state-of-the-art AR models, end-to-end speedup rises from 1.02-1.08x at 480P to 1.62x at 720P and 1.95x at 1080P, entirely without training, while keeping VBench quality within 0.2 points of full attention and largely flicker-free. A single threshold set transfers unchanged across four backbones, three resolutions and both clip lengths (5 s and 30 s), indicating that the taxonomy reflects model-independent structure rather than per-checkpoint calibration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.