acceptodds
Under review as a conference paper at ICLR 2027

Focus-KV: Head-Wise KV Cache Compression Focus On Reasoning Heads

Abstract

Large language models rely on long chain-of-thought (CoT) to solve complex reasoning problems, but the long decoding trajectories make the key–value (KV) cache a major memory cost. Existing KV-cache compression methods reduce this cost by evicting tokens or adaptively reallocating cache budgets at runtime. However, they either assign the same budget to all attention heads, which gives substantial memory to heads that contribute little to reasoning, or rely on runtime attention statistics that can be biased by the already-compressed cache. To address this, we propose Focus-KV, a static head-wise KV-cache compression algorithm that allocates more budget to attention heads that are important for reasoning, and the per-head budgets remain unchanged at runtime. Experiments across multiple models and benchmarks show that Focus-KV consistently outperforms strong baselines, particularly under highly constrained budgets. We further demonstrate that the identified reasoning-critical heads capture reasoning ability that is general rather than confined to specific tasks: the importance signal obtained from a mathematical probe set can be directly applied to reasoning tasks in other domains without re-estimation. Focus-KV exploits this cross-domain generality to better preserve the model’s reasoning ability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.