acceptodds
Under review as a conference paper at ICLR 2027

MEKV: Controlling Concentration in Query-Agnostic Visual KV Compression

Abstract

Large vision-language models require compressing visual KV caches for efficient multi-query serving. When compression happens before user queries are known, a natural approach is to retain tokens according to a relevance signal—yet we show that high-quality signals can degrade accuracy when over-concentrated, starving unseen questions of needed evidence. We develop a concentration theory along the maximum-entropy temperature path, proving that a surrogate accuracy objective is single-peaked with the optimum given by a moment-matching condition. This yields MEKV, a simple, training-free method that separates signal quality from concentration control via a single temperature parameter. The theory further predicts that the right strategy is model-specific: models whose optimum requires strong tempering need concentration control first, while those near the raw signal benefit more from fine-grained token selection. On full validation across three benchmarks, MEKV improves over coverage baselines by up to 17 pp, with the temperature transferring across tasks without re-tuning. The prediction is borne out by a strategy reversal: same-signal per-head top- falls far behind MEKV on Qwen yet excels on InternVL, and placing per-head selection inside calibrated quotas repairs the failure while preserving the gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.