VQKV: High-Fidelity KV Cache Compression for Memory-Limited Long-Context Inference
Abstract
A compact key–value (KV) cache that represents every token reduces memory requirements for long-context inference. We present VQKV, which combines representation learning, output-aware calibration, and fused attention using shared post-RoPE codebooks. With residual SimVQ and asymmetric key/value capacity, VQKV retains 98.4% of the full-cache LongBench average on LLaMA3.1-8B at 82.8% code-length compression. It achieves the highest aggregate scores among the evaluated baselines at 90% and 95% compression. With the quantizer fixed, static value calibration improves 32K RULER by 3.22 points in the primary paired evaluation while preserving LongBench quality. It uses 128 KiB of parameters and leaves the encoder and cache format unchanged. Matched iterative and analytic MSE controls show the benefit of optimizing model outputs in addition to reconstruction. Fused FlashAttention integrates the correction into value lookup and computes attention directly from compact codes. By trading encoding and lookup computation for cache capacity, VQKV supports frozen-model inference on evaluated workloads whose full-cache memory requirements exceed GPU capacity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.