CommitStream: Can Models Keep Their Answers Correct in Streaming Video?
Abstract
Advances in embodied intelligence and world modeling place increasing demands on models’ video understanding. For video assistants, a key challenge is that an answer that is correct when issued can become wrong as the scene changes. If an assistant neither updates nor withdraws it, users may continue acting on outdated information. Reliable assistance therefore requires maintaining an answer that is supported by the available evidence throughout an ongoing task. We introduce CommitStream, a benchmark that jointly evaluates when models should wait for evidence, provide an answer, and revise it as events unfold. To support this evaluation, a reusable pipeline generates and refines annotations of when answers become supported and when they must change, followed by a final human review. The resulting benchmark contains 1,006 tasks from 307 video episodes across six domains. Evaluation follows causal real-time playback, using streaming inference for models with a streaming input interface and asynchronous sliding-window inference for those without one. Under the official ±1 s tolerance, the best model with streaming inference remains correct throughout 19.88% of tasks; the best with asynchronous sliding-window inference reaches 29.03%. Even the latter leaves stale or wrong answers in effect for 13.94% of scored time. These results identify sustained answer reliability as an unresolved challenge and establish CommitStream as a testbed for studying when video assistants should answer, wait, and update. Code and data are available at https://anonymous.4open.science/r/commitstream.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.