VideoLive: Benchmarking Stateful Real-Time Audio-Visual Agents
Abstract
Multimodal large language models (MLLMs) have demonstrated strong video understanding capabilities, yet they are mostly evaluated on answers given at isolated moments. A correct answer at one moment does not show whether a model can keep taking correct and timely interactive actions as the scene evolves. Real-time interaction therefore requires a model to generate answers and to decide when to **ANSWER**, when to **WAIT**, and when to **REPAIR** an answer that has become stale. We introduce VideoLive, a *stateful streaming* benchmark that evaluates real-time audio-visual interaction through a unified causal protocol. VideoLive continuously requires an agentic MLLM to make decisions from the accumulated evidence available up to the current moment, autonomously choosing whether to answer, wait, or repair, and evaluates its capability across the full lifecycle of an answer, from forming an initial answer through retention to update. We evaluate 29 model configurations across four lifecycle types: Retention, Replacement, Withdrawal, and Formation. The results show that forming a correct initial answer does not guarantee timely subsequent updates: the best configuration completes **23.50%** of answer lifecycles, and across all 29 configurations only **21.11%** of valid initial answers were revised correctly and in time after the state changed. Even when an independent probe identifies the new activity, **44.26%** of the public answers from incremental interfaces remain incorrect two seconds later, revealing a substantial gap between content understanding and continuous interactive decision-making. We further propose **SPOT**, a post-training method that guides MLLMs to learn when to publish, retain, or update an answer by supervising their replies conditioned on the displayed answer; it improves MMDuet2 on every VideoLive metric, raising LCS from 3.25 to 8.50 and Overall from 15.44 to 25.72, and improves response timing on three held-out streaming benchmarks. We hope VideoLive helps move research toward reliable real-time interaction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.