Semantically Aligned KV Cache Compression for Efficient Agentic Autoregressive Video Generation
Abstract
Agentic autoregressive video generation offers a promising approach to controllable video synthesis through semantic planning, sequential generation, and visual verification. However, growing visual histories enlarge the key-value (KV) cache, increasing memory use and attention costs across candidate rollouts. KV cache compression reduces this overhead, but existing methods primarily optimize individual generators without explicitly accounting for the planner's semantic constraints. Applying them directly can assign low importance to states associated with particular requested attributes or actions, potentially weakening semantic fidelity. In this paper, we introduce **SemanticKV**, a training-free method for semantically aligned KV cache compression. To our knowledge, **SemanticKV** is the first KV cache compression method for agentic autoregressive video generation that explicitly uses planner-extracted semantic constraints to guide retention. Text cross-attention identifies constraint-related video positions, whose self-attention queries are collected into separate groups. Attention weights are averaged within each group and maximized across groups, giving each cached entry a high score when it is strongly used by any represented constraint. A single top- selection retains high-scoring entries under the memory budget, leaving their keys, values, and positions unchanged. For fixed query groups and equal entry costs, this selection minimizes an additive upper bound on worst-group discarded attention. The method operates independently within each candidate without retraining or additional video generations. Our evaluation spans four AR video backbones and three benchmarks, covering 5-, 30-, and 60-second generation; our method achieves strong performance across these settings. On Causal Forcing at a compression ratio of , **SemanticKV** obtains a VBench Total Score of , compared with for the strongest compressed baseline and for Full KV. Meanwhile, it achieves a speedup of over Full KV and reduces KV-cache memory from GB to GB, corresponding to a reduction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.