SuBitKV: Sub-Bit KV Cache Compression for Autoregressive Video Generation
Abstract
Autoregressive video generation relies on a growing key–value (KV) cache to retain information from previous frames, which raises a memory challenge as video length increases. However, retaining the full KV cache at high precision becomes impractical under limited GPU memory. Discarding past states or reducing their precision saves memory but can compromise consistency in appearance, structure, and motion over time. To address this challenge, we introduce SuBitKV, a training-free KV cache compression framework that supports storage below one bit per original KV element. Its key idea is to preserve a low-memory approximation of the retained KV cache as a persistent base, then allocate additional bits across the KV cache within the memory budget to improve reconstruction fidelity. We construct this persistent base using residual vector quantization, which jointly represents the elements of each key or value vector with a small set of compact indices rather than storing each element separately. As the cache grows, SuBitKV reduces these extra bits to stay within budget while retaining the persistent base. Experiments on three video-generation and world-model backbones show a strong memory–quality trade-off. On Wan2.2, SuBitKV achieves compression with near-BF16 VBench scores and reaches under tighter memory budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.