acceptodds
Under review as a conference paper at ICLR 2027

SVI-Bench: Evaluating Human-Perceived Interaction Trajectories in Streaming Video Systems

Abstract

Users do not experience a streaming vision-language assistant as a collection of isolated responses. They experience a continuous interaction trajectory in which the assistant watches a changing scene, decides when to speak or stay silent, reacts with some delay, and remembers information over time. The quality of this interaction therefore depends not only on what the assistant says, but also on how its behavior is aligned with the evolving visual context. To better evaluate these interaction demands, we introduce SVI-Bench, a trajectory-centric benchmark for streaming vision-language interaction. SVI-Bench evaluates five user-observable aspects: proactive triggering, silence correctness, latency, response correctness, and delegation with long-horizon memory. It contains 75 expert-authored items across nine realistic interaction categories, with 272 anchored dimension-level criteria. Central to SVI-Bench is InteractFlow, a trace-based evaluation framework that synchronizes model outputs with the visual states visible when they occur. This turns temporal interaction from abstract metadata into directly inspectable multimodal evidence and allows responses to be judged in the visual context in which they actually occurred. It also preserves each system's streaming behavior and supports traceable rescoring. Across five deployed configurations, JoyAI-VL-Interaction achieves the highest Overall score. The results also show that current systems are much better at staying silent than at speaking when needed. SVI-Bench therefore shifts streaming evaluation from scoring individual responses to evaluating the complete interaction experienced by the user. All videos, item-specific scoring criteria, recordings, and InteractFlow will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.