acceptodds
Under review as a conference paper at ICLR 2027

RateQuant: Attention-Aware Rate–Distortion Allocation for KV Cache Quantization

Abstract

Effective KV cache compression combines accurate quantization with precision allocation to preserve model quality within a memory budget. Allocation is coupled through existing attention error and inputs from earlier quantized layers. We introduce RateQuant, a rate–distortion framework allocating KV precision across layers and heads through measured attention responses. These responses capture how adjacent bit-width changes reinforce or cancel the existing attention error. Dynamic programming combines them into budget-preserving joint updates, and remeasurement refreshes the responses for the next round. Full-model cross-entropy on separate calibration data then selects the final allocation. Our analysis connects normalized attention distortion to task loss under bounded suffix sensitivity and establishes exact optimization of the local additive objective. Across three models and two quantizers, RateQuant achieves the lowest perplexity in all 18 settings spanning 2–3-bit budgets, with reductions of 12.6–94.5% at 2 bits and 3.2–10.3% at 2.5 bits versus each setting's strongest tested baseline. Frozen allocations achieve the highest three-model mean scores in all 30 LongBench evaluation settings (five tasks, two quantizers, and three budgets), without task-specific recalibration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.