acceptodds
Under review as a conference paper at ICLR 2027

EOV-Bench: An Evidence-Oriented Benchmark for Long-Horizon Streaming Video Understanding

Abstract

Streaming video benchmarks are typically organized around predefined tasks, offering limited control over the evidence structure of questions. We introduce EOV-Bench, an evidence-oriented benchmark that constructs questions bottom-up from reusable, multi-scale video annotations. Evidence units are selected and composed before question formulation, controlling question type, temporal span, evidence distribution, and compositional demand—factors that shape difficulty. Reusing the annotation pool makes fuller use of long videos and yields diverse, explicitly grounded questions spanning current perception, historical retention, and cross-temporal reasoning. EOV-Bench contains 4,337 open-ended questions and supports matched retention analysis under a strict causal streaming protocol. Across evaluated strategies, high-fidelity recent frames favor current perception, whereas textual memory better preserves historical information over long time horizons. Ablations further show that long-term retention depends substantially on how observations are described and organized, not only on retrieval. These findings motivate AsyncMemFlow, a training-free framework that decouples memory construction from latency-sensitive answering. An independently configurable pipeline generates detailed descriptions and asynchronously consolidates them into structured memory. At query time, the answerer adaptively uses recent frames or the latest available memory, supplemented with observations not yet incorporated into it. This enables stronger memory construction without blocking online answering. Experiments demonstrate improved historical and joint reasoning while preserving strong current perception, with further gains on OVO-Bench.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.