acceptodds
Under review as a conference paper at ICLR 2027

What to Forget: Contribution-Aware Memory for Streaming Video Understanding

Abstract

Multimodal Large Language Models (MLLMs) have made impressive strides in offline video understanding tasks. However, adapting them to streaming settings is far from straightforward, as the ever-growing visual context quickly exhausts GPU memory. While prior work mitigates this by compressing the cache and retaining tokens with high attention scores, we find that these scores do not always reflect a token's contribution to the model output. In this paper, we formalize token contribution through the Leave-one-out (LOO) reference, defined as the shift in model output distribution caused by removing a single token. Utilizing the self-prediction objective, we derive the first-order KV Taylor score that shows stronger agreement with LOO than attention-based scores. Building upon this, we thus propose CALM, i.e., Contribution-Aware Layer-wise Memory, a training-free caching strategy that allocates retention budgets based on KV Taylor scores and applies conditional novelty reranking to suppress redundant tokens. To retain long-range context, we further maintain a lightweight text memory that stores compact summaries of past video chunks. Experiments on StreamingBench and OVO-Bench demonstrate that CALM achieves state-of-the-art streaming video understanding results, with gains of 6.25% on StreamingBench and 11.82% on OVO-Bench (12.77% on the Real-Time and 10.87% on the Backward Tracing).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.