EviStream: Evidence-Preserving Memory for Streaming Video Understanding
Abstract
Streaming video assistants must continuously process an unbounded visual stream before knowing what users may ask in the future, requiring efficient memory without losing evidence needed for later reasoning. Existing approaches mainly reduce visual computation or build retrievable memories, but often leave unclear which visual evidence should survive compression and how it should be recovered when needed. We introduce EviStream, a training-free evidence-preserving memory framework based on MLLMs. EviStream follows the principle of preserving residual information rather than repeated observations through complementary semantic memory and visual evidence. Incremental Semantic Memory (ISM) organizes the incoming stream into adaptive windows, removes temporal redundancy, and incrementally consolidates each window into semantic deltas, yielding a compact evolving textual history. To complement this semantic abstraction, Memory-Aligned Evidence Keyframe Selection (MEKS) preserves sparse visual evidence according to temporal novelty, the information gap left by semantic consolidation, and query relevance when available. Once a query arrives, Continual Evidence-Polling Reasoning (CEPR) retrieves relevant semantic memory and complementary visual evidence together with live observations, while continuously consolidating and re-retrieving new observations when current evidence is insufficient. EviStream achieves state-of-the-art real-time understanding performance on StreamingBench and OVO-Bench, reaching 84.1% and 76.8% accuracy, respectively, without additional training. Code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.