acceptodds
Under review as a conference paper at ICLR 2027

Calibration Timing and Model Sensitivity in Rotated KV-Cache Quantization

Abstract

Does a rotated uniform KV-cache quantizer require locally fitted scales? A comparison can conflate scale granularity with calibration timing and metadata cost. We audit these factors with non-overlapping finite windows, a dense non-evicted reference, and one-write quantize–dequantize simulation. Following development experiments, a frozen 80-condition comparison crosses model, dataset, rotation seed, calibration timing, granularity, and symmetry, with eight-bit controls. On Llama-3.1-8B-Instruct, four-bit per-write global calibration yields post-prefill perplexities of 7.508–7.557 on WikiText-2 test (reference 7.404), whereas prefill-frozen global configurations yield 125.414–39,255.244. The timing effect persists on SQuAD contexts and for two new rotations. Grouped per-write scales reduce error further but incur more metadata, preventing an equal-total-budget claim. Qwen2.5-7B-Instruct behaves differently: all tested four-bit configurations fail severely across both datasets and rotations, while eight-bit controls remain close to their dense references. Unquantized cache and rotation-roundtrip controls also remain close on a diagnostic window. These results support calibration- and model-specific conclusions rather than universal necessity of local scales. They do not establish packed-kernel speed, cross-domain generalization, or a unique mechanism for the Qwen low-bit failure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.