acceptodds
Under review as a conference paper at ICLR 2027

Elastic-QVG: Efficient Long Video Generation with Elastic KV-Cache Quantization

Abstract

Auto-regressive video generation uses a growing KV-cache to retain long-range context. Existing KV-cache quantization reduces the memory cost per token but still allows total memory usage to grow with context length, whereas eviction bounds memory by discarding cached context entirely. To bridge this gap, we present **Elastic-QVG, a training-free framework that progressively lowers cache precision without re-quantizing quantized caches.** A bit-level sliding window extends conventional token-level windowing to individual bit-planes, keeping recent chunks at higher precision while truncating older chunks to fit a fixed memory budget. Each chunk is quantized once, and its precision can subsequently decrease as new context arrives or the available memory budget shrinks. A truncation-exact quantizer enables hardware-friendly truncation directly on the quantized KV-cache, with no additional error relative to direct lower-bit quantization. Its nested quantization grids make dropping bit-planes exact, while separately allocated bit-plane storage and fused dequantization turn this property into memory savings without repacking the cache. Across LongCat-Video, HY-WorldPlay, and LingBot-World, Elastic-QVG improves fidelity over static QVG and retains VBench scores close to BF16 at 13–15× compression. It reduces KV memory by up to 12.7× with less than 5% latency overhead on H100, and avoids KV offloading on an RTX 5090 to achieve a 4.3× speedup over offloaded BF16-cache inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.