FactorKV: A Marginal Coverage Selector for Low-Budget KV Cache Compression in LVLMs
Abstract
Large vision-language models (LVLMs) rely on growing key-value (KV) caches for autoregressive generation, making cache compression important for long visual contexts. Existing methods improve how limited capacity is allocated across attention heads and how individual KV pairs are scored. Yet under tight budgets, top-k selection can retain tokens from answer-relevant regions while still producing answers with missing or incorrect words. Static ranking can favor overlapping content, leaving complementary answer details insufficiently represented. We propose FactorKV, a KV cache selector based on utility-weighted marginal coverage of the full compressible history in key space. Given a host compressor's budget, protected positions, and utility scores, FactorKV first constructs a pool of high-scoring candidates. It then greedily selects the candidate that provides the greatest additional coverage and updates the represented content after each selection. Historical positions already well covered contribute less to subsequent gains, allowing candidates that represent complementary details to gain priority. Each retention decision therefore accounts for both the utility of the historical content and the coverage supplied by the current selected set. FactorKV integrates with existing compressors without requiring retraining or increasing the cache budget. Experiments show gains across multiple cache allocators and LVLMs, particularly under tight budgets. FactorKV repairs some omission and substitution errors made by the original methods, while others remain unresolved. Under low-budget, decoding latency remains comparable to the base methods, and total latency remains lower than FullKV despite additional prefill computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.