ViFlow-KV: Visual-Flow Guided KV Cache Eviction For Multimodal Reasoning
Abstract
Long-horizon multimodal reasoning combines dense visual inputs with growing textual trajectories, making KV cache memory a major inference bottleneck for Large Multimodal Models (LMMs). KV cache eviction mitigates this burden by selectively retaining high-utility states under a limited cache budget. However, existing methods typically apply a modality-agnostic eviction policy, overlooking both the distinct retention requirements of visual and textual states and their evolving interaction during reasoning. Our attention analysis characterizes a heterogeneous visual flow: LMMs gradually shift the access to visual information from the original visual inputs toward later generated hidden states. Furthermore, later queries revisit both the original visual evidence and earlier generated states that carry visual information, revealing a long-range reuse of visual information. Guided by these observations, we propose ViFlow-KV, a visual-flow-guided KV eviction framework for multimodal reasoning. At each eviction step, ViFlow-KV assigns modality-specific budgets to visual and textual states, while incorporating visual-anchor signals to preserve reasoning-relevant visual evidence across evolving carriers. By preserving visual evidence across evolving visual and textual carriers, ViFlow-KV maintains critical context throughout long reasoning while keeping cache highly compressed. Extensive experiments show that ViFlow-KV consistently outperforms existing eviction methods, matching full attention performance at 2k token budget while achieving up to 3.05 throughput speedup.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.