VimPro: Long Video Understanding via Proxy-Guided Agentic Reasoning over Hierarchical Memory
Abstract
Long-video understanding remains challenging due to lengthy temporal contexts and sparsely distributed question-relevant evidence. Although memory-augmented methods organize long videos into structured memories to preserve and retrieve long-range information, they still face two major challenges. First, they use a static memory as the sole inference context, limiting their ability to adaptively refine or supplement question-specific evidence. Second, similarity-based retrieval may return semantically related yet answer-irrelevant or misleading memories, whereas LLM-based agentic retrieval incurs high inference cost and is difficult to control. To overcome these limitations, we propose VimPro, a proxy-guided framework that constructs and dynamically maintains a hierarchical memory for long-video understanding. During inference, the answer agent reasons over the evolving memory and requests supplementary evidence when needed, while a lightweight proxy, MemLens, filters noisy candidates into compact answer-supporting evidence. MemLens is trained through cold-start supervised fine-tuning followed by answer-gated reinforcement learning with tailored, fine-grained rewards. Extensive experiments demonstrate that VimPro consistently outperforms both closed-source and open-source SOTA methods. All datasets, code, and trained models will be publicly released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.