acceptodds
Under review as a conference paper at ICLR 2027

Query-adaptive Visual-semantic Evidence Allocation for Streaming Video Understanding

Abstract

Streaming video understanding requires answering questions over an ever-growing visual history under limited memory and inference budgets. Existing methods typically manage this history through fixed-capacity memory compression or similarity-based frame retrieval, but largely rely on fixed compression or retrieval criteria without explicitly modeling the specific historical evidence requirements of each query. This overlooks a key property of streaming questions: their historical evidence requirements are inherently heterogeneous. To address this issue, we propose a training-free framework for query-adaptive visual-semantic evidence allocation that jointly determines whether to access history, which temporal forms of visual evidence to prioritize, and which long-range semantic evidence to retrieve. Given a query and recent observations, the framework jointly predicts whether historical access is needed and the query’s preferences over static states, dynamic context, and recency-sensitive evidence. Questions supported by recent obser- vations are answered directly through Fast–Slow routing. For queries requiring history, the same evidence preferences guide both fine-grained visual selection and event-level semantic ranking. The visual branch combines query-specific preferences for static, dynamic, and recency-sensitive evidence to select a compact and diverse set of historical frames. The semantic branch propagates the preference signals to events in a hierarchical caption memory, fuses the event scores with the same query preferences, and complements preference-guided ranking with VLM- based caption retrieval. The resulting visual and semantic evidence is integrated with recent observations for final answering. Our framework achieves 67.79% accuracy on OVO-Bench and 77.01% across the 14 visual-only tasks of Streaming- Bench, outperforming the compared methods on both benchmarks. These results demonstrate the value of coordinating history access and visual-semantic evidence selection according to query-specific requirements.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.