acceptodds
Under review as a conference paper at ICLR 2027

ProjCache: Output-Projected KV Cache Eviction with OBS-Inspired Compensation

Abstract

Evicting key-value (KV) pairs reduces the memory cost of large language model inference but perturbs attention outputs. We introduce ProjCache, which unifies KV cache eviction and compensation under a shared layer-wise objective: the reconstruction error measured after the attention output projection. We parameterize both operations using token masks applied to attention weights and extend the saliency analysis of Optimal Brain Damage (OBD) to these mask coordinates. By summing all terms in the Taylor expansion of the reconstruction error, we derive an exact single-pair eviction score. This score accounts for attention renormalization and retains cross terms between query heads sharing a KV head in grouped-query attention. For compensation, we adapt Optimal Brain Surgeon (OBS) to jointly optimize the masks of retained tokens using a quadratic approximation of the same reconstruction objective. Experiments on LongBench and RULER with LLaMA-3.1-8B and Qwen3-8B show that replacing heuristic attention scores with ProjCache scores improves performance in 106 of the 108 evaluated settings, with gains of up to 3.4 points on LongBench and 25.6 points on RULER. ProjCache achieves performance comparable to state-of-the-art methods. For compensation, ProjCache lowers the reconstruction error by 17% to 44% in all 24 LongBench and RULER-4K settings, whereas prior compensation methods increase it in every setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.