acceptodds
Under review as a conference paper at ICLR 2027

RAPID: Extreme KV Cache Compression At Near-Full Quality from the Context Alone

Abstract

The size of KV caches is a dominant bottleneck in long-context model serving. Modern KV compression algorithms fall short on at least one of the following objectives. 1. To maintain uncompressed generation quality. 2. To compress fast, without heavy training, data generation, or model use (beyond context itself). 3. To achieve best possible compression: SOTA as of mid-2026 is over 10× rate as compared to full KV cache storage at half-precision floating points. We introduce RAPID algorithm (Rank And Precision allocation with Independent per-layer Distillation) which our evaluations show to be the closest to satisfying all three objectives. RAPID combines layer-wise latent representations, cross layer budget allocation, per token dimension reduction, per dimension quantization, and intra-layer optimization using prefill query vectors only. RAPID doesn't need training data and access to the model beyond each layer's attention projections. A variant of our algorithm, RAPID++, is always within 2.5 points of the full cache quality at 20× compression, on every evaluated benchmark and model, including Qwen3-4B, a 117B-parameter mixture-of-experts GPT-OSS-120B, and an MLA model whose cache is already compressed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.