Elastic-QVG: Efficient Long Video Generation with Elastic KV-Cache Quantization
Abstract
Auto-regressive video generation uses a growing KV-cache to retain long-range context. Existing KV-cache quantization reduces the memory cost per token but still allows total memory usage to grow with context length, whereas eviction bounds memory by discarding cached context entirely. To bridge this gap, we present **Elastic-QVG, a training-free framework that progressively lowers cache precision without re-quantizing quantized caches.** A bit-level sliding window extends conventional token-level windowing to individual bit-planes, keeping recent chunks at higher precision while truncating older chunks to fit a fixed memory budget. Each chunk is quantized once, and its precision can subsequently decrease as new context arrives or the available memory budget shrinks. A truncation-exact quantizer enables hardware-friendly truncation directly on the quantized KV-cache, with no additional error relative to direct lower-bit quantization. Its nested quantization grids make dropping bit-planes exact, while separately allocated bit-plane storage and fused dequantization turn this property into memory savings without repacking the cache. Across LongCat-Video, HY-WorldPlay, and LingBot-World, Elastic-QVG improves fidelity over static QVG and retains VBench scores close to BF16 at 13–15× compression. It reduces KV memory by up to 12.7× with less than 5% latency overhead on H100, and avoids KV offloading on an RTX 5090 to achieve a 4.3× speedup over offloaded BF16-cache inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.