StreamingMEM-R1: Learning memory evolution for streaming video understanding
Abstract
Streaming video understanding requires models to interpret incoming content causally, preserve information that may disappear from view, and respond at the appropriate time. Existing approaches typically retain fragmented historical representations and rely on predefined memory-update rules, limiting their ability to track evolving events and coordinate evidence retention with response decisions. To address these challenges, we propose **StreamingMEM-R1**, a framework that jointly learns memory evolution, evidence routing, and response timing. At each streaming step, the model performs structured memory evolution through two complementary modules: stable memory and dynamic memory. The evolved states parameterize a response policy that estimates whether an answer is warranted. For response steps, a routing variable specifies stable, dynamic, or joint memory as the answer context. We construct 100K supervised fine-tuning examples and 50K reinforcement-learning examples to teach these decisions explicitly. Streaming-aware fine-tuning preserves causal visual and semantic context, while memory-oriented reinforcement learning provides verifiable supervision for memory evolution and trajectory-level correctness. Experiments across four benchmarks demonstrate that StreamingMEM-R1 enables effective memory evolution, accurate long-horizon reasoning, and efficient streaming video understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.