Q-Span: Query-Adaptive Look-Back Windows for Streaming Video Understanding
Abstract
What a streaming video assistant can answer is bounded by how much history it keeps, but decided by which slice of that history the model is actually shown, and that slice cannot be fixed. Measuring accuracy against the look-back window on three streaming benchmarks, we find the window a question needs varies at least in length, and that populations disagree about it as sharply *within* a benchmark as between them: on one level of OVO-S-Bench a s window beats a s one by points, and on another the longer window wins by . No single window serves a stream of mixed questions. We propose Q-Span, which reads the question and nothing else to choose the span. Which of three windows a question calls for, a close look at the present, a recent span, or a thin pass over the whole prefix, is settled by its wording, so one text call picks the window before a frame is decoded. How far back the evidence sits is not settled by the wording: the same sentence can sit three seconds from its answer in one recording and three minutes in another. That residue is read off the frames, in one further call on frames the answer will use anyway. The method is training-free and its rules name no benchmark's templates, though written for the kinds of question streaming benchmarks ask: planner, probe and answering model are one frozen backbone under different prompts. Across three backbones and three benchmarks it holds both ends at once, which no fixed window does, and the gain follows the mixture: three points where a benchmark spreads its questions, and a tie where they do not.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.