Towards a Dynamic and Fixed-budget Memory Bank for Streaming Video Understanding
Abstract
Currently, streaming video understanding is still a daunting task for existing Multimodal Large Language Models (MLLMs). Its difficulties not only lie in handling the ever-increasing video frames, but also in the unpredictability of future video content and input instructions. In this paper, we study this task from the perspective of constructing a dynamic but fixed-budget memory bank, and propose a novel and training-free approach termed CausalMem. CausalMem targets at constructing a dynamic visual memory update mechanism, thereby maximizing the amount of information in streaming video within a limited memory space, much like human brain. In practice, CausalMem estimates the redundancy of visual tokens and updates the memory bank via carefully constructed online semantic basis, which models the principal semantics of video streams. To validate CausalMem, we apply it to three representative MLLMs, namely LLaVA-OneVision, Qwen2.5-VL and Qwen3-VL respectively, and conduct extensive experiments on both streaming and offlineEvaluated under streaming protocol. video understanding benchmarks. The experimental results not only show the great advantages than existing methods using streaming or offline settings, e.g., and average accuracy gains respectively, but also witness the superior semantic preservation for streaming videos, e.g., using 12 token budgets to memorize hour-long streaming videos, which achieves more than 20 visual token compression ratio and only occupies about 82 MB storage. Our code is given in the supplementary materials.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.