acceptodds
Under review as a conference paper at ICLR 2027

Gated Attention-aware Layout-Preserving Physical KV Cache Compression in Vision-Language Models

Abstract

Large vision-language models cache visual representations in decoder key-value (KV) states throughout autoregressive generation, creating a substantial memory burden. However, existing KV compression methods mainly target language-style, one-dimensional sequences and overlook the two-dimensional spatial organization of visual states, so aggressive compression can discard the spatial evidence the decoder still relies on. To address this issue, we present Gated Attention-aware Layout-Preserving KV compression (GALP-KV), a training-free method that physically compresses the decoder KV cache after multimodal prefill while preserving model-aligned visual layout. GALP-KV identifies the visual span in the full decoder KV cache, maps visual entries to the model-aligned grid, protects all non-visual states, and scores entries by target-model attention utility augmented with a weak geometric prior. Its primary causal mechanism is a budget-gated, layer- and KV-head-specific selection policy: a shared layout-preserving stable route at moderate budgets, and a per-head proxy route under aggressive compression. Experimental results show that the geometric prior contributes no measurable marginal effect once uniform and tail anchors are present. On Qwen2.5-VL, GALP-KV reduces physical KV memory by 48.76% while preserving near-Full-KV question-answering (QA) behavior, and remains stable at a 56.23% reduction, where the layout-aware mIoU advantage over attention-only selection widens. On LLaVA-1.5-7B, the same policy achieves a 52.60% reduction and changes only 224 predictions relative to Full KV, fewer than uniform, random, and SnapKV retention. Matched-count analysis shows that GALP-KV leaves the smallest largest-empty-radius (spatial blind region) of all compared selectors at the main budget under a paired bootstrap, forming a top tier with NACL across budgets. Grounding remains more sensitive than QA, and the current prototype does not yet provide robust end-to-end latency acceleration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.