acceptodds
Under review as a conference paper at ICLR 2027

Memory-Efficient Autoregressive Video Generation via Hierarchical Mixed-Precision KV Cache Quantization

Abstract

Autoregressive video generation models face a critical KV cache memory bottleneck as the demand for long video synthesis grows. Existing LLM quantization methods perform poorly when applied to video generation models due to the high dynamic range of video KV cache activations and their heightened sensitivity to quantization errors. We propose HiMix, a hierarchical mixed-precision KV cache quantization framework for memory-efficient autoregressive video generation. Specifically, HiMix first determines layer-wise bit budgets offline based on per-layer sparsity analysis, then performs runtime channel-aware quantization guided by Q-K saliency: important channels receive higher bit-widths (4-bit) while less important ones are quantized to 2-bit or 1-bit, with spatiotemporal Hadamard smoothing applied prior to quantization to reduce error. Experiments on the Self-Forcing and Causal Forcing models demonstrate that HiMix achieves nearly 7.1x compression ratio for KV cache while preserving video generation quality and long-range consistency comparable to the BF16 baseline, substantially reducing memory consumption.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.