RateQuant: Attention-Aware Rate–Distortion Allocation for KV Cache Quantization
Abstract
Effective KV cache compression combines accurate quantization with precision allocation to preserve model quality within a memory budget. Allocation is coupled through existing attention error and inputs from earlier quantized layers. We introduce RateQuant, a rate–distortion framework allocating KV precision across layers and heads through measured attention responses. These responses capture how adjacent bit-width changes reinforce or cancel the existing attention error. Dynamic programming combines them into budget-preserving joint updates, and remeasurement refreshes the responses for the next round. Full-model cross-entropy on separate calibration data then selects the final allocation. Our analysis connects normalized attention distortion to task loss under bounded suffix sensitivity and establishes exact optimization of the local additive objective. Across three models and two quantizers, RateQuant achieves the lowest perplexity in all 18 settings spanning 2–3-bit budgets, with reductions of 12.6–94.5% at 2 bits and 3.2–10.3% at 2.5 bits versus each setting's strongest tested baseline. Frozen allocations achieve the highest three-model mean scores in all 30 LongBench evaluation settings (five tasks, two quantizers, and three budgets), without task-specific recalibration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.