DualStream: Episodic Feedback Shapes Visual Memory for Long-Horizon Streaming Video Understanding
Abstract
Vision-language models commonly construct video context offline from a complete video. Streaming video understanding instead requires context to be built under resource constraints as frames arrive, often before the user question is known. Many streaming methods update this context based on visual redundancy or local changes, but visual novelty alone does not determine what is worth retaining. The importance of incoming details also depends on earlier events and can change as the episode unfolds. To address this gap, we present DualStream, a training-free framework that uses evolving textual episodic memory to guide visual retention. The visual stream maintains a compact token pool, which the textual stream periodically reads to update a running summary of the observed video. Each update also produces a future focus statement that guides the ranking of visual tokens from subsequent frames. Newly retained evidence then informs the next episodic update, forming a closed loop that adapts visual selection as the video develops. On the long subset of Video-MME, DualStream reaches 60.0% accuracy, exceeding the strongest online baseline by 4.7 percentage points while dropping 86.1% of visual tokens. Beyond these gains, at the same visual token budget, DualStream outperforms both visual memory alone and combined visual and textual memory without feedback.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.