acceptodds
Under review as a conference paper at ICLR 2027

SPARE: Shared Perception and Adaptive Retrieval of Evidence for Streaming Video Understanding

Abstract

Streaming video understanding requires answering questions from the observed video, with evidence drawn from recent observations or earlier events. Historical context can improve recall but may interfere with judgments based on recent frames. Its usefulness depends on both when it is introduced and which clips are selected, while relevance alone does not guarantee sufficient evidence for an answer. We propose SPARE, a framework that combines adaptive historical retrieval with visual encoder sharing. A lightweight gate predicts whether recent observations provide sufficient evidence to answer the question. When additional evidence is predicted to be necessary, a multimodal embedding model retrieves relevant historical clips to supplement the current context. LoRA instruction tuning trains the understanding model to use recent observations and retrieved history jointly and abstain when the evidence remains insufficient. To reduce redundant visual parameters, the understanding and retrieval branches share a visual encoder while retaining separate language backbones. Four linear adapters align the shared features with the retrieval model through distillation at the visual token and final retrieval embedding levels. With visual-only inputs, SPARE achieves 77.78% on the evaluated Real-time and Backward task groups of OVO-Bench and 73.31% on StreamingBench. Compared with Qwen3-VL-8B-Instruct evaluated at 2 fps, SPARE improves accuracy by 20.68 percentage points on OVO-Bench and 14.15 percentage points on StreamingBench.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.