Think Global, Act Local: Efficient KV Cache Quantization via Global Prefill Calibration of Early Decoding
Abstract
Efficient mixed-precision KV cache quantization requires identifying which cached tokens will remain important during future decoding, before aggressively compressing them. To address this challenge, we propose ForeKV, an efficient KV cache quantization framework that estimates future token importance from early decoding queries and calibrates this local estimate with global query information from prefill. Our key insight is that the first few decoding queries provide the earliest direct evidence of the model's emerging attention behavior, but such evidence is limited by the short observation window. Therefore, we delay compression for a few decoding steps to collect real decoding queries, estimate token importance from these local observations, and calibrate the estimate using the mean query over the global prefill sequence. The resulting query directly scores cached keys to guide mixed-precision allocation without materializing the full attention matrix. We further design a paged attention kernel for K caches with token-wise K4/K2 allocation and channel-wise quantization, enabling efficient decoding directly over the compressed cache. Experiments across multiple LLM backbones on LongBench show that ForeKV remains close to BF16 at an effective compressed-cache storage ratio of approximately , with the average score dropping by no more than approximately 1%. On Qwen3-8B, our end-to-end ForeKV implementation achieves up to the generation throughput of the BF16 baseline under the same 80 GB GPU memory budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.