acceptodds
Under review as a conference paper at ICLR 2027

EviStream: Evidence-Aligned Audio-Visual Decisions over Partially Observable Video Streams

Abstract

Vision-language models (VLMs) could enable assistants that continuously monitor video streams and respond when a user-specified condition is met. Yet recognizing the current scene is insufficient for deciding when to respond: some events depend on earlier observations, while spoken evidence becomes available only after transcription. Under bounded context, a streaming assistant must preserve relevant task progress and determine whether the available evidence supports a response. We introduce EviStream, a training-free framework for evidence-aligned audio-visual decisions over partially observable video streams. Its central idea is to manage observed evidence, persistent task progress, and temporary model judgments separately. A bounded KV cache retains recent visual and transcript evidence, while a temporary inference branch produces compact, evidence-indexed observations. A persistent state tracks established prerequisites and consumed observations to determine whether to wait or trigger a notification. Incremental speech transcripts are incorporated as they become available, and answers are generated from the frozen triggering context. Using frozen Qwen3-VL-32B-Instruct, EviStream achieves a joint timing-and-content score of 25.3 on all 774 StreamArena Pro questions, compared with 4.7 for direct window monitoring with the same backbone. Removing persistent state, audio, or sliding KV reduces the score to 10.3, 22.1, and 14.9, respectively. These results support the effectiveness of coordinating evidence retention and task progress for proactive streaming responses without task-specific fine-tuning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.