RAPID: Extreme KV Cache Compression At Near-Full Quality from the Context Alone
Abstract
The size of KV caches is a dominant bottleneck in long-context model serving. Modern KV compression algorithms fall short on at least one of the following objectives. 1. To maintain uncompressed generation quality. 2. To compress fast, without heavy training, data generation, or model use (beyond context itself). 3. To achieve best possible compression: SOTA as of mid-2026 is over 10× rate as compared to full KV cache storage at half-precision floating points. We introduce RAPID algorithm (Rank And Precision allocation with Independent per-layer Distillation) which our evaluations show to be the closest to satisfying all three objectives. RAPID combines layer-wise latent representations, cross layer budget allocation, per token dimension reduction, per dimension quantization, and intra-layer optimization using prefill query vectors only. RAPID doesn't need training data and access to the model beyond each layer's attention projections. A variant of our algorithm, RAPID++, is always within 2.5 points of the full cache quality at 20× compression, on every evaluated benchmark and model, including Qwen3-4B, a 117B-parameter mixture-of-experts GPT-OSS-120B, and an MLA model whose cache is already compressed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.