acceptodds
Under review as a conference paper at ICLR 2027

ReSight: Reconstructing Fine-Grained Visual Evidence for Streaming Video Understanding

Abstract

Understanding a changing world requires reasoning about evidence that is no longer visible. Streaming video understanding demands this ability under bounded memory, yet observations arrive before their relevance to future questions is known. We introduce ReSight, which reconstructs fine-grained visual evidence from a compact persistent state. Our key insight is to give memory a stable visual reconstruction target while learning separately how recovered evidence should influence an answer. ReSight maintains a fixed-size visual field through recursive least-squares updates in fixed temporal and appearance coordinates. At query time, a lightweight reader reconstructs spatially resolved features and pools them into a bounded visual context, deferring question-dependent interpretation to the decoder. To account for imperfect reconstruction, asymmetric attention preserves a perception reference within a shared decoder, and a learned signed gate regulates the deviation proposed by history-informed reasoning. This separates preserving observations from deciding their influence, without replaying past frames or optimizing parameters at query time. Trained on only 10% of StreamTTT's training data, ReSight surpasses StreamTTT on seven of eleven task metrics across OVO-Bench and StreamingBench and matches it on one. It improves action sequence identification by 10.1 percentage points and sequential question answering by 7.20 points, while raising real-time visual understanding accuracy by 2.76 points. Moreover, ReSight reduces per-frame ingestion time by 19% and persistent state storage by 57% relative to StreamTTT. These results highlight the effectiveness of reconstructive memory with higher data efficiency and inference efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.