acceptodds
Under review as a conference paper at ICLR 2027

Positional Specialization for Autoregressive Video KV-Cache Compression

Abstract

KV-cache compression for autoregressive video is usually framed as deciding how much of the past to preserve. Using persistent, position-matched interventions, we show instead that different temporal positions require different kinds of information. Distant states, including pinned and mid-lag positions, can be spatially collapsed from 1,560 tokens to a single token with negligible loss, whereas the same collapse on recent states causes severe frame-to-frame flicker. Spatial compressibility, however, does not imply dispensability: evicting distant states degrades image quality and nearly halves motion, whereas retaining each as a single token does not, and replacing them with another video's cache shifts output semantics increasingly with semantic distance. These findings motivate AxisKV, a two-axis allocation that spatially compresses distant states while quantizing recent states. At cache compression on Self-Forcing, AxisKV outperforms matched-memory uniform pooling and, against uniform quantization, trades a loss in aesthetic quality for higher image quality; it also remains usable at , where uniform quantization has no valid operating point. The positional specialization transfers across related autoregressive video generators.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.