CoCoKV: Consensus- and Completeness-Aware KV Cache Compression for LLM Inference
Abstract
Efficient long-context inference in Large Language Models (LLMs) is severely constrained by the memory footprint of the Key-Value (KV) cache. Recent prompt-stage methods reduce this overhead by estimating historical-token importance from observation-window attention and retaining only a compact subset of the KV cache. However, their reliance on head-specific token importance gives rise to two potential limitations: i) Head-wise Attention Partiality, where individual attention heads can capture specialized yet partial views of token importance, and ii) Token-wise Information Omission, where selecting highly attended tokens leads to semantic omission while overlooking other crucial information. To address these limitations, we present CoCoKV, a novel KV cache compression framework designed to retain both reliable and comprehensive evidence under constrained cache budgets. Specifically, CoCoKV introduces a soft cross-head consensus mechanism that uses collective signals to refine head-specific attention while preserving head distinctiveness. Building on the calibrated importance scores, a cross-token completeness selection strategy is applied to promote broader semantic coverage. Evaluations across Needle-in-a-Haystack and LongBench datasets on three LLMs demonstrate that CoCoKV outperforms competitive baselines, validating its effectiveness and robustness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.