Quantize, Evict, or Both? When Mixed KV-Cache Allocation Pays, and Why its Distortion Proxy Misleads
Abstract
A key-cache budget can be spent by quantizing every token, by evicting tokens, or by mixing zero and finite precisions across tokens; values stay exact. Hybrid systems exist. We ask when this mixed interior beats both endpoints, and whether an answer measured on attention-output distortion carries over to end tasks. If eviction is treated as a zero-rate tier, a finite-bit tier is dominated when its logitnoise cost exceeds the cost of zero rate. The share of heads whose 2-bit tier is dominated orders the interior’s exactly recomputed output-error advantage across six LLMs and 25 model–context cells (Spearman ρ = −0.978; −0.94 controlling for model). Where the interior stops paying, however, depends on the architecture, and longer contexts move a model toward eviction. Over 4,096 decode tokens a head’s best policy family is stable, but its token allocation goes stale within a few steps. On retrieval tasks we compare against three dense quantizers (TurboQuant, KIVI, KVQuant) and six eviction methods (H2O, SnapKV, DropKV, Ada-KV, LaProx, OBCache). Routed per KV head, the allocator that is optimal under a local surrogate of output distortion beats every eviction method (mean score 0.77 vs. 0.60), partly by sending 1–25% of heads to dense quantization; its mixed interior alone scores 0.47. Yet dense quantization at the same key code bits does better on all three end-task models, and the best quantizer depends on the model. TurboQuant wins on Llama-3.1-8B (0.99 vs. 0.65 for the allocator) and Mistral-7B (0.95 vs. 0.70). On Qwen3-8B TurboQuant loses to the allocator (8k–32k: 0.87 vs. 0.96), and per-channel KIVI or KVQuant win (0.99). Output distortion ranks methods only weakly and overstates the cost of rotated quantization noise. The allocator mainly fails by truncating multi-token answers; smoothing its token scores over neighboring positions largely removes this failure (Llama, 8k–32k: about 40% to 10% of answers). At both levels, the choice of baselines and of when the cache is compressed moves verdicts more than prompt sampling does.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.