acceptodds
Under review as a conference paper at ICLR 2027

KVCapsule: Efficient Sequential KV Compression for VLMs with Asymmetric Redundancy

Abstract

Visual tokens can dominate the key-value (KV) cache of long-context vision-language models (VLMs), increasing persistent memory and decoding traffic. Many visual sequence-reduction methods prune or merge positions using importance estimates obtained before or early in decoding, making discarded KV states unavailable if their relevance changes later. We show that visual KV exhibits sequence-level redundancy, decoding-dependent attention, and distinct redundancy structures for keys and values. Based on these observations, we introduce KVCapsule, an asymmetric visual-KV compression framework that reduces persistent cache storage while maintaining an approximate full visual-position representation for subsequent attention. KVCapsule reconstructs visual keys from selectively retained positions and compresses visual values using sequence-level PCA, while keeping the pretrained VLM backbone frozen. Across VLM backbones and image and video benchmarks, KVCapsule achieves a strong quality-memory trade-off. At 10K context and batch size 1, it reduces the measured persistent KV-related footprint by 55.0% and improves decoding throughput by 24.0%; at larger-batch long-context settings, throughput improves by up to , while our analytical memory model predicts up to a smaller persistent footprint beyond the break-even regime.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.