KVRecon: KV Cache Merging via Reconstruction Error Minimization
Abstract
Large language models (LLMs) increasingly rely on long contexts, making the memory footprint of key-value (KV) caches a major bottleneck for efficient inference and motivating the need for effective KV cache compression. Prior KV cache merging methods rely on predefined criteria for selecting entries to merge, or heuristic merging operations to combine them, which can introduce substantial discrepancies in model predictions after compression. In this paper, we introduce KVRecon, a KV cache merging method that minimizes the reconstruction error in the output of the attention module. It groups KV entries into multiple sets, and derives merged key and value representations for each set to preserve the output after merging. We further identify KV entries whose attention logits or value representations differ substantially from those of their corresponding merged entries to prevent inaccurate merging. Extensive experiments across multiple LLMs and downstream tasks demonstrate that KVRecon consistently outperforms existing KV cache compression methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.