acceptodds
Under review as a conference paper at ICLR 2027

Videopus: Parallel Multimodal Evidence Seeking for Long-Horizon Video Understanding

Abstract

Long-horizon video understanding requires models to locate sparse evidence across extended timelines while preserving the visual detail needed for accurate reasoning. Existing agentic systems often inspect candidate intervals sequentially, repeatedly processing an interaction history. Moreover, local observations are usually summarized in text, making it difficult for the global reasoner to verify fine-grained visual details or connect evidence across intervals. We introduce Videopus, a multi-agent framework in which an Orchestrator coordinates interval-specific Executors. It has two components. Parallel Task Spawn groups inspections derived from the same evidence state into a concurrently executable batch. Multimodal Evidence Aggregation returns structured local findings with a compact set of supporting visuals selected from each Executor's input. The Orchestrator uses this evidence to resolve ambiguous descriptions and compare objects or event stages across the video. Local perception remains interval-specific, while global reasoning operates over a compact, visually inspectable evidence state. Videopus alternates parallel inspection with evidence-guided refinement. Across six video benchmarks, Videopus consistently outperforms the state-of-the-art methods while reducing token use and sequential rounds and increasing persistent-context reuse across reasoning turns.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.