GOPAgen: Codec-Aware Agentic Long-Video Understanding with Structured Memory
Abstract
Agentic long-video question answering often relies on sparse RGB sampling and caption retrieval. This reliance can lead to missed brief motion events and repeated processing of irrelevant intervals. We introduce GOPAgen, a framework that integrates codec-native Groups of Pictures (GOPs) and motion vectors into an agentic retrieval pipeline. GOPAgen constructs global textual memory and uses two-stage query-conditioned temporal selection. A global-sufficiency check precedes local evidence extraction. Caption and motion agents represent selected intervals as typed memory pages containing complementary appearance and motion evidence, temporal metadata, and provenance. A GOP-Tree organizes these pages for conditional retrieval when the global evidence is insufficient. GOPAgen achieves 65.3% accuracy on MotionBench Test and 78.7% on EgoSchema, exceeding the reported agentic baselines listed for these benchmarks. It also achieves 73.2% on LongVideoBench validation and 77.3% on MLVU, while remaining below the strongest listed baseline on LVBench and Video-MME Long. These results support codec-aware structured memory as an effective interface for selective long-video reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.