What to Forget: Contribution-Aware Memory for Streaming Video Understanding
Abstract
Multimodal Large Language Models (MLLMs) have made impressive strides in offline video understanding tasks. However, adapting them to streaming settings is far from straightforward, as the ever-growing visual context quickly exhausts GPU memory. While prior work mitigates this by compressing the cache and retaining tokens with high attention scores, we find that these scores do not always reflect a token's contribution to the model output. In this paper, we formalize token contribution through the Leave-one-out (LOO) reference, defined as the shift in model output distribution caused by removing a single token. Utilizing the self-prediction objective, we derive the first-order KV Taylor score that shows stronger agreement with LOO than attention-based scores. Building upon this, we thus propose CALM, i.e., Contribution-Aware Layer-wise Memory, a training-free caching strategy that allocates retention budgets based on KV Taylor scores and applies conditional novelty reranking to suppress redundant tokens. To retain long-range context, we further maintain a lightweight text memory that stores compact summaries of past video chunks. Experiments on StreamingBench and OVO-Bench demonstrate that CALM achieves state-of-the-art streaming video understanding results, with gains of 6.25% on StreamingBench and 11.82% on OVO-Bench (12.77% on the Real-Time and 10.87% on the Backward Tracing).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.