KVarN: Variance-Normalized KV-Cache Quantization for Improved Reasoning in LLMs
Abstract
Test-time scaling improves language-model reasoning but increases KV-cache memory during long-horizon generation. Existing KV-cache quantization methods are commonly evaluated in static prefill settings, which do not capture feedback from repeatedly attending to a quantized cache. We show that this feedback increases attention-output error over time and identify token-scale distortion as a major contributor to the largest key-quantization errors. We introduce KVarN, a calibration-free, linear KV-Cache quantization method to effectively suppress token scale errors and error accumulation. At an effective storage cost of 2.25 bits per element, KVarN achieves the best overall quality against strong quantization baselines on AIME 2024, MATH-500, HumanEval, and IFEval. It also reduces attention-output error under our pseudo-decode diagnostic. In end-to-end vLLM measurements across three models, KVarN improves decode throughput by 10–28% over the 16-bit KV-cache baseline; at a nominal 3-bit precision, it achieves 1.44–2.00× the throughput of the vLLM TurboQuant implementation. Code for the KVarN method is given in the supplementary.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.