acceptodds
Under review as a conference paper at ICLR 2027

KVarN: Variance-Normalized KV-Cache Quantization for Improved Reasoning in LLMs

Abstract

Test-time scaling improves language-model reasoning but increases KV-cache memory during long-horizon generation. Existing KV-cache quantization methods are commonly evaluated in static prefill settings, which do not capture feedback from repeatedly attending to a quantized cache. We show that this feedback increases attention-output error over time and identify token-scale distortion as a major contributor to the largest key-quantization errors. We introduce KVarN, a calibration-free, linear KV-Cache quantization method to effectively suppress token scale errors and error accumulation. At an effective storage cost of 2.25 bits per element, KVarN achieves the best overall quality against strong quantization baselines on AIME 2024, MATH-500, HumanEval, and IFEval. It also reduces attention-output error under our pseudo-decode diagnostic. In end-to-end vLLM measurements across three models, KVarN improves decode throughput by 10–28% over the 16-bit KV-cache baseline; at a nominal 3-bit precision, it achieves 1.44–2.00× the throughput of the vLLM TurboQuant implementation. Code for the KVarN method is given in the supplementary.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.