acceptodds
Under review as a conference paper at ICLR 2027

AutoStream: Automatic Context Expansion for Streaming Video Understanding

Abstract

Online video understanding requires models to perceive ongoing events while retaining relevant information from the past. Existing training-free methods, however, struggle to balance these two demands: the accumulated memory maintained by history-centric approaches may interfere with real-time perception, whereas recent-window approaches preserve real-time responsiveness at the cost of discarding long-term evidence. We present AutoStream, a training-free framework that realizes a real-time perception first, history on demand principle through automatic context expansion. During stream ingestion, Query-Ready Multi-Context Cache (QRMC) prepares hierarchical caches that span real-time perception and comprehensive historical context. Novelty-Weighted Reservoir Sampling (NWRS) retains temporally dispersed and informative historical evidence, while time-adaptive capacity reallocation progressively shifts the context budget toward long-range coverage as the stream grows. At query time, Confidence-Guided Context Expansion (CGCE) progressively exposes broader history based on model confidence, while its History Shortcut directly activates the comprehensive history for queries identified as history-dependent. By avoiding extra retrieval and auxiliary models, AutoStream maintains low query-time latency. With Qwen2.5-VL-7B as the backbone, AutoStream improves performance from 73.3% to 81.0% on StreamingBench and from 44.7% to 58.2% on the Backward subset of OVO-Bench. With Qwen3-VL-8B, it achieves 67.8% on OVO-Bench and 82.3% on StreamingBench.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.