acceptodds
Under review as a conference paper at ICLR 2027

Unified Damage-Based KV Cache Compression for Vision-Language Models

Abstract

Vision-Language Models (VLMs) cache the keys and values of context tokens for autoregressive decoding, but the KV cache grows linearly with context length and can be dominated by thousands of visual tokens. Existing eviction methods typically rank tokens by accumulated attention, measuring the sum of attention weights a token receives across queries rather than what is lost when it is removed, while separate heuristic sets per-layer budgets. We present DEKV (Damage-Equalized KV Eviction), a training-free framework that uses a single damage metric to determine both token importance and per-layer budgets. The metric is the exact attention-output perturbation caused by evicting one token, obtained in closed form from a softmax leave-one-out identity. DEKV sums these per-token damages as an additive surrogate for joint eviction, which we bound analytically and validate empirically. Normalizing by each layer's output scale makes damages comparable across layers, and a single global threshold, equivalent to water-filling over the per-layer damage curves, allocates the memory budget without an external allocation policy. DEKV scores only text-query attention rows and keeps the cache at exactly the budget after prefill and after every decoding step. Across six VLMs from the LLaVA, Qwen3-VL, and Phi-3.5 families on four VQA benchmarks, DEKV outperforms the strongest evaluated baseline, TGV-KV, in 22 of 24 settings at a 5% KV budget, retaining 86.6% of full-cache accuracy on average versus 80.0%, with gains of up to 27.5 points on ChartQA and 13.0 points on DocVQA for Phi-3.5-vision, while reducing KV cache memory by 20×.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.