Query-Agnostic Visual KV Compression at Matched Memory Budgets
Abstract
Prefix caching allows vision-language models (VLMs) to reuse the key-value (KV) cache of an image across questions, which requires compression to be completed before the question is known. Most existing visual KV compression methods are query-aware and therefore cannot be applied in this scenario; for the query-agnostic setting, there is a lack of a systematic and fair comparison under the same memory allocation. In this paper, we investigate query-agnostic visual KV compression under matched memory budgets, across two VLM families and four visual question answering benchmarks. On our fixed paired evaluation manifests, retaining deterministically spaced evicted tokens has higher mean scores than merged tokens in all eight model–task comparisons and matches or exceeds tiled moment summaries in all comparisons. Besides, pre-registered falsification tests and an exploratory mixture study do not establish a summary advantage. Meanwhile, we provide an attention-mass bound on the maximum contribution of an eviction summary and analyze three mechanisms that limit summary methods: dilution of individual value outliers, instability of the dispersion correction, and per-block RoPE interactions in merged keys. Finally, quantization behavior depends on the grouping axis and on metadata cost. Together, these controlled comparisons clarify how byte allocation, retained-token fidelity, and task structure affect query-agnostic visual KV compression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.