Fund Regions, Not Tokens: Query-Conditioned Budget Allocation for Extreme KV-Cache Eviction in Vision-Language Models
Abstract
Vision-language models (VLMs) need thousands of visual tokens for multi-page documents and high-resolution charts. The per-request key–value (KV) cache limits how many requests a device can serve concurrently. Existing eviction methods from text LLMs keep the global top-k tokens of a one-dimensional sequence. Visual evidence, however, occupies contiguous two-dimensional regions, and a region stays usable only if enough of its tokens remain. At extreme budgets of 1–3%, a global ranking spreads the cache over many locations and leaves each region too few tokens. To this end, we propose CARVE, which treats KV eviction as query-conditioned budget allocation to regions. CARVE partitions visual tokens into layout-aware cells, commits most of the budget to a few query-ranked regions, and selects tokens within them. When selection-time diagnostics find this allocation unsuitable, CARVE falls back to attention-based selection, and its gated mode CARVE-G to conservative token-level eviction. Experiments cover four models and four benchmarks at equal KV budget. On DocVQA at a 1% budget, CARVE retains 74.4% of the full-cache score on InternVL3-8B, where the strongest text-guided baseline retains 40%. CARVE-G stays within 1.6 points of its text-guided fallback in all sixteen model–budget settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.