acceptodds
Under review as a conference paper at ICLR 2027

Before the Question Arrives: Causal Evidence Budgeting for Streaming Video Understanding

Abstract

Streaming video understanding is crucial for AI systems to interpret dynamic environments, yet providing accurate, real-time responses remains challenging. Models must process a growing visual history within limited context and memory budgets. Existing methods reduce this pressure through visual-token compression or key–value (KV) cache management, but they often fix what to retain at frame arrival, before considering earlier and later evidence together. Such early decisions can discard evidence needed by future questions, and later retrieval cannot recover it. To address this mismatch, we introduce StreamingRETAKE, a training-free framework for causal evidence budgeting that allocates visual-token and KV capacity over time. Visual-token capacity is directed toward informative frames to build a compact evidence pool, from which shared evidence is selected jointly after more frames arrive. Then retrieval combines evidence importance with question relevance to build the answer context. The framework bounds visual-history storage and answer contexts as the stream grows. Extensive experiments show that our proposed method achieves state-of-the-art accuracy of 75.52% and 81.80% on OVO-Bench Real-Time and StreamingBench real-Time, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.