When Per-Head KV Precision Helps: Interactions, Allocation, and Scale Timing
Abstract
A common approach to mixed-precision KV caches allocates bits per attention head using isolated sensitivity scores: quantize one head, measure the loss increase, and assign more bits to the heads whose quantization increases loss most. We test three assumptions behind this recipe with paired, teacher-forced measurements on Qwen3-0.6B, Qwen3-1.7B, and Mistral-7B. First, joint low-bit damage is superadditive: when every head is quantized to two bits under a symmetric integer code, joint damage is about 3.3 times the sum of single-head damages on Qwen3-1.7B and 3.0 times on Mistral-7B, whereas four-bit damage on Qwen3-1.7B is nearly additive. Second, under a min-max code, whether keys or values are more sensitive depends on when quantization scales are computed. When scales use complete token groups, two-bit values cause much greater loss than two-bit keys; when scales use only entries already written, this value-side penalty disappears on all three models and the ordering reverses on both Qwen models. Third, the best allocator depends on the operating point: at about 3.3 average bits, measured selection on Qwen3-1.7B lowers loss by 0.044 nats relative to an equal-cost static first-layer rule, whereas at 2.75 bits a static first-16-layer rule gives the lowest held-out loss on Mistral-7B. In a serially filling cache, two-bit keys with four-bit values improve on uniform three bits by 0.16 nats on Mistral-7B at equal modeled cost. Per-head scores thus support useful allocation when combined with joint evaluation of the assembled cache, an explicit scale rule, and an equal-cost static baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.