VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents
Abstract
Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from *semantic thrashing*: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it has already found. First, we provide a structural argument showing that append-only memory can incorporate newly observed target evidence but cannot remove accumulated noise or prevent the growth of ordered context without a rewrite operator. Second, motivated by this analysis, we propose VideoLoop, a multimodal agent with two coupled loops. The outer loop reasons over the video and the inner loop, after each step, retrieves artifacts from an unbounded filesystem of past observations and intermediate analysis, and rewrites a bounded working memory. Extensive experiments demonstrate the effectiveness of VideoLoop, which improves four popular LVLM backbones in a plug-and-play manner, with an average gain of 4.2% points over baseline on VideoMME (*long*). Further analysis of working memory suggests that VideoLoop mitigates semantic thrashing, achieving 81% context-only judge accuracy on a human-reviewed challenging subset versus 59% for the append-only agent. Notably, with Gemini 3.1 Pro, our VideoLoop achieves new state-of-the-art results on all three benchmarks: 88.3% on VideoMME (long), 88.8% on VideoMMMU, and 80.9% on LongVideoBench (long).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.