acceptodds
Under review as a conference paper at ICLR 2027

StreamAct: Streaming or Just Answering? Diagnosing Streaming Intelligence Illusions in Video-Language Models

Abstract

Are video-language models truly capable of streaming video understanding, or do they merely answer questions from a given causal context? Many streaming-like evaluations provide a causal prefix or recent window and score the resulting answer at an evaluator-chosen moment. This can make models appear streaming-capable without testing whether they autonomously maintain relevant state, judge evidence sufficiency, or choose appropriate actions as the stream unfolds. We call this mismatch the Streaming Intelligence Illusion. To diagnose it, we introduce StreamAct, which evaluates streaming video understanding as a strict-causal online trajectory over state, readiness, action, and answer. StreamAct-Bench evaluates 5,436 replayable trajectories under a strict-causal protocol. Target-time answerability alone does not establish this behavior: Qwen3-Omni-30B-A3B achieves 45.02% target-time answer accuracy, yet only 6.68% of its first self-timed responses are timely and correct. To test whether this missing evidence-to-action conversion is learnable, StreamAct-SFT provides explicit supervision over state, readiness, and action, selectively improving these trajectory-level behaviors beyond answer-only and format-only training. These results establish the distinction between streaming and just answering: Streaming means maintaining state and choosing the right action as evidence unfolds.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.