PACE-VL: Preserving Attention via Clustered Evictions for Vision-Language Models
Abstract
Vision-language models that ingest high-resolution images or multi-page documents produce long, predominantly visual key–value (KV) caches, making decode-time memory and latency primary bottlenecks. Under a strict cache budget, the central challenge is preserving task quality when the vast majority of physical slots must be aggressively evicted. We introduce PACE-VL, a training-free KV-cache compression method that jointly designs retention, residual aggregation, and decode-time reading. Specifically, query states from the question or instruction tokens score the history to retain the highest-scoring entries, called evidence tokens, with their original keys and values. The remaining historical entries are clustered into anchors, cache entries that each store a selected original key, an aggregated value sum, and a multiplicity count—the number of tokens represented. At decode time, attention incorporates cluster multiplicities into normalization, so each cluster maintains its aggregate contribution under a shared-anchor-key approximation. Evidence tokens, anchors, and a short, uncompressed recent window occupy a cache of exactly the budgeted size. Across five visual-heavy benchmarks on three open VLMs spanning the Qwen and Llama-3 lineages, a single fixed configuration of PACE-VL without any fine-tuning outperforms all compared compression baselines in all 21 evaluated settings at a 10% cache budget. It matches or exceeds full-cache decoding in 11 settings and otherwise trails it by at most 1.87 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.