acceptodds
Under review as a conference paper at ICLR 2027

RRC-KV: Range-Driven Rotation and Centering for Two-Bit KV Caches

Abstract

Scalar INT2 KV caches have only four reconstruction levels per group, making accuracy sensitive to quantization coordinates. We present RRC-KV, which learns per-head key and per-layer value rotations through range-based objectives. A query-based auxiliary regularizer limits coordinate sensitivity concentration, while static key centering shifts the encoding origin with its contribution restored in attention logits. The method requires neither downstream labels nor quantizer backpropagation and supports mixed-precision attention with fused Triton kernels. Across four models and five reasoning and coding benchmarks, RRC-KV achieves a macro-average score of 79.54 versus 80.04 for BF16. Under the same KV storage layout, it outperforms the spectral-rotation baseline OSCAR in every tested RULER setting across two models and 8K–64K contexts (87.19 versus 81.61 on average), while a long-context gap to BF16 remains. Across three Qwen models, scores on two coding tasks vary by at most 3.05 points between two calibration configurations, compared with up to 34.69 points for OSCAR on Qwen3-32B. At 64K tokens, the KV representation provides 6.91× compression relative to BF16, including metadata and retained BF16 windows.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.