Do Streaming Video Models Remember The Past? Verifying and Evolving Memory System for Videos
Abstract
Memory management is a fundamental capability for long-horizon video understanding, where models must store and retrieve useful information from a growing stream of observations. Yet, it remains unclear whether existing benchmarks actually require the memory, as even a simple memory-free system can outperform dedicated memory-augmented models on standard streaming-video benchmarks, raising questions about what these evaluations truly capture. We introduce MemProof, a benchmark that verifies memory dependence by ensuring each question is unsolvable from language or recent observations alone, but becomes solvable once the necessary past evidence is revealed. With memory dependence explicitly verified, we next ask whether the memory-management strategy itself can be improved without updating any model parameters. We introduce EvoHarness, which keeps all neural components frozen and evolves modular programs for organizing, retrieving, and presenting memory through teacher-guided failure diagnosis and paired validation. Experiments show that strong performance on conventional streaming benchmarks can be achieved even without memory, whereas MemProof requires models to actually use the past: under the same backbone, evolution improves the base harness by +8.0 points, and the memory organization and packing mechanisms of EvoHarness transfer to other benchmarks. The code/data will be made available online.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.