EviStream: Evidence-Aligned Audio-Visual Decisions over Partially Observable Video Streams
Abstract
Vision-language models (VLMs) could enable assistants that continuously monitor video streams and respond when a user-specified condition is met. Yet recognizing the current scene is insufficient for deciding when to respond: some events depend on earlier observations, while spoken evidence becomes available only after transcription. Under bounded context, a streaming assistant must preserve relevant task progress and determine whether the available evidence supports a response. We introduce EviStream, a training-free framework for evidence-aligned audio-visual decisions over partially observable video streams. Its central idea is to manage observed evidence, persistent task progress, and temporary model judgments separately. A bounded KV cache retains recent visual and transcript evidence, while a temporary inference branch produces compact, evidence-indexed observations. A persistent state tracks established prerequisites and consumed observations to determine whether to wait or trigger a notification. Incremental speech transcripts are incorporated as they become available, and answers are generated from the frozen triggering context. Using frozen Qwen3-VL-32B-Instruct, EviStream achieves a joint timing-and-content score of 25.3 on all 774 StreamArena Pro questions, compared with 4.7 for direct window monitoring with the same backbone. Removing persistent state, audio, or sliding KV reduces the score to 10.3, 22.1, and 14.9, respectively. These results support the effectiveness of coordinating evidence retention and task progress for proactive streaming responses without task-specific fine-tuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.