acceptodds
Under review as a conference paper at ICLR 2027

GeoKV: Geometry-Aware KV Cache Compression via Global Structural Contribution and Key Rectification

Abstract

Vision-Language Models (VLMs) have demonstrated remarkable capabilities, yet the rapidly growing memory consumption of Key-Value (KV) caches poses a major challenge for efficient inference. Existing KV cache compression methods predominantly characterize token importance from a query-dependent functional perspective, such as attention or local relevance, while overlooking the global structural role of cached representations. Consequently, tokens with low current-query relevance may still be important for preserving underrepresented directions and the global geometry of the cached-key space. To address this limitation, we propose GeoKV, a global-structure-aware framework for KV cache compression. Specifically, we introduce Global Structural Contribution (GSC), a query-independent criterion derived from the marginal gain of a regularized log-determinant objective, and perform greedy selection to preserve structurally valuable tokens. We further introduce a lightweight geometry-aware key rectification module based on a low-rank feature-space transformation to mitigate compression-induced geometric distortion. Extensive experiments on VLMs demonstrate that GeoKV effectively reduces KV cache size while maintaining model performance. Comprehensive analyses further show that structural contribution captures complementary information beyond query-dependent functional relevance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.