Block-GTQ: RoPE-Block Rate Allocation for KV-Cache Quantization
Abstract
Larger models and longer contexts make the KV cache a memory and bandwidth bottleneck. We introduce Block-GTQ, a framework that uses the two-dimensional frequency blocks of rotary position embeddings (RoPE) to allocate cached-key precision. A blockwise logit-error bound motivates a label-free query–key energy score, which we combine with codec-dependent rate penalties. Under equal-cost budgets and diminishing marginal gains, largest-marginal-gain allocation exactly minimizes the separable integer proxy. Our TurboQuant-MSE (TQ-MSE) instantiation groups equal-rate blocks and supports attention over packed caches without materializing a decoded fp16 KV cache. Across ten model/path configurations, K-only diagnostics reduce layer-averaged logit mean absolute error by about – relative to uniform TQ-MSE, improving all layers at both 2- and 3-bit average logical key rates. At K2V2 (2 average logical bits per K/V coordinate) on Llama-3.1-8B-Instruct, Block-GTQ raises the six-task needle-in-a-haystack (NIAH) average from to and LongBench-EN from to ; K2V2 gains also hold on the tested 32B and 70B models. On DeepSeek-R1-Distill-Qwen-7B at K3V2, without fp16 sink or recent-token buffers, it achieves avg@8 on AIME 2024/2025 versus fp16’s , where uniform TQ-MSE scores zero on both. On Qwen2.5-3B-Instruct, packed K3V3 serving on one H800 reduces resident KV footprint by a factor of and retains perplexity close to fp16 at 128K, while uniform TQ-MSE degrades severely. At a 1M-token context, Block-GTQ reduces decode latency by 8.0% relative to FP16 FlashAttention-2.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.