Redundancy-Aware Visual KV Memory for Streaming Video Understanding
Abstract
Streaming video understanding requires MLLMs to maintain informative visual context under a bounded memory budget as frames arrive continuously. Existing visual KV compression methods largely evaluate KVs as individual units, although feature-similar KVs can encode highly redundant visual content. Such redundancy affects both cross-layer capacity allocation and within-layer KV retention: repeated representations can distort token-level statistics used to estimate layer-wise memory demand and consume multiple retention slots while contributing limited additional information. This work introduces RaKV, a training-free framework for redundancy-aware visual KV memory management. RaKV defines a unified KV score that combines query-agnostic importance with feature representativeness and uses it for both memory decisions. For cross-layer allocation, we aggregate feature-similar KVs into visual groups and measure group-level entropy, yielding a layer-wise demand estimate that is less sensitive to repeated representations. For within-layer retention, the same redundancy-aware score favors KVs that are important and representative of distinct visual patterns. Experiments across streaming video benchmarks and MLLM backbones show that RaKV consistently improves streaming video understanding under a bounded visual KV budget, with gains of up to 9.50% over the original backbone. Our code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.