Learning Memory Update Policies for Streaming Video Understanding through Reinforcement Learning
Abstract
Streaming video understanding requires compact memory that preserves useful information as new observations arrive. Existing training-free methods often compress visual memory using heuristic criteria, while conventional KV-cache pruning has limited applicability to the linear-attention layers of hybrid backbones. We propose MURL(Memory Update Reinforcement Learning), a framework that learns visual-token retention policies for recursive memory updates through reinforcement learning. With the vision-language backbone frozen, a lightweight selector scores tokens from retained history and newly arrived observations. A history gate combines current contextual scores with stored importance estimates to guide repeated compression. We train the policy using video-level caption feedback from a frozen backbone. Experiments with Qwen3.5-4B and Qwen3.5-9B on StreamingBench, OVO-Bench and Video-MME under a streaming protocol demonstrate the effectiveness of learned memory updates for streaming video understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.