What Does Full Attention Add? State-Complementary Visual KV Compression for Hybrid Vision-Language Models
Abstract
Vision-language models are increasingly adopting hybrid architectures that interleave full-attention layers with recurrent modules. Recurrent modules summarize historical information in fixed-capacity states, while full attention preserves direct access to historical visual tokens through explicit key–value (KV) caches. Existing visual KV compression methods typically assess entry importance through visual contributions to full attention, without explicitly accounting for information already represented in recurrent states. This motivates us to consider the additional information that full attention provides beyond recurrent-state representations when compressing visual KV entries in hybrid architectures. We propose Full-Attention Extra Semantic Information (FA-ESI), a training-free method for active visual KV compression that explicitly models this complementarity. For a target full-attention layer, FA-ESI replays the preceding Gated DeltaNet on the target layer's normalized inputs to construct an auxiliary recurrent state. It estimates full-attention extra information by comparing FA values with the corresponding state readouts. Current attention weights condition this information on the query. Projection onto the subspace of top-ranked prediction candidates followed by alignment scoring yields token-level FA-ESI scores for dynamic visual KV selection during decoding. Across benchmarks spanning multi-image understanding, long-video understanding, and GUI-agent tasks, FA-ESI achieves approximately 95% compression of active visual KV entries while retaining 98.984%–99.243% of Vanilla's performance on average.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.