HelixKV: Allocating Bits Where They Matter for Autoregressive Video Diffusion
Abstract
Autoregressive video diffusion models enable streaming and long-horizon video generation, but their key–value (KV) caches grow linearly with sequence length, creating an increasingly severe memory bottleneck. Existing KV-cache quantization methods typically rely on uniform or coarse-grained precision allocation. This wastes memory on quantization-tolerant cache regions and leaves fewer bits for regions where precision matters most. To alleviate this problem, we introduce HelixKV, a training-free KV-cache compression framework that allocates bits where they matter. Specifically, HelixKV characterizes the impact of KV quantization through model-output distortion and allocates precision along two complementary dimensions. Across attention heads, it assigns head-specific bit-widths to keys and values according to their quantization sensitivity. Along the temporal axis, it further differentiates the precision of recent and historical KV states, allowing each head to retain only the precision required by different parts of its history. Combining these distortion profiles with byte-accurate storage costs, HelixKV solves a global budgeted allocation problem to jointly allocate KV precision across the head and temporal axes. By compressing quantization-tolerant cache regions and reallocating the saved memory to regions where precision matters more, HelixKV improves generation quality at a fixed KV-cache budget while preserving the complete KV history. Experiments on Self-Forcing-1.3B and HY-WorldPlay-5B show that HelixKV achieves and KV-cache compression, respectively, while retaining 99.9% of the VBench Total score on both models, demonstrating a favorable memory–quality trade-off.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.