Beyond Linear Surrogates: Softmax-Aware Information Capacity for KV Cache Eviction
Abstract
Key-value (KV) cache eviction is essential for scaling long-context inference in Large Language Models (LLMs). However, existing policies predominantly rely on empirical heuristics, lacking a principled characterization of token utility under the inherently nonlinear softmax attention mechanism. In this work, we revisit KV cache eviction through the lens of local information geometry, modeling attention as a nonlinear Gaussian communication channel. By locally linearizing the attention mapping through a first-order Taylor expansion, we derive the Jacobian Information Capacity, an information-theoretic objective that jointly captures query relevance, softmax sensitivity, and structural diversity. Guided by this formulation, we introduce Jacap, a capacity-aware eviction method that combines softmax-aware importance weighting with statistical leverage scores for subset selection. Extensive experiments across diverse model architectures and benchmarks show that Jacap strong and consistent improvements over existing eviction methods across most settings, particularly under high compression ratios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.