acceptodds
Under review as a conference paper at ICLR 2027

StreamWeave: Curating Multi-Timescale Visual Evidence for Streaming Video Understanding

Abstract

Multimodal large language models have advanced offline video understanding, but streaming video systems must answer questions using only observations available when each question arrives. As the observed prefix grows beyond the model's finite visual context, different questions make different observations relevant. A fixed frame selection cannot reliably preserve all such evidence. Even when past frames are retained, event descriptions written before a question arrives may not reveal which frames contain the visual evidence it seeks. We introduce StreamWeave, a training-free agent that curates multi-timescale visual evidence for each question. As video arrives, it builds a searchable visual history linked to original frames. At query time, a recent glance and a broader look-back provide the first-call evidence; the MLLM answers or requests additional source frames through history retrieval. Across three streaming video understanding benchmarks and five frozen MLLMs ranging from an 8B dense model to a 397B mixture-of-experts model, StreamWeave achieves competitive results against published streaming methods. Extensive ablation studies show the effectiveness of its components, temporal layout, and retrieval choices.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.