ASVBench: Asynchronous Response to Fine-Grained Triggers in Streaming Videos
Abstract
Streaming video understanding is a pivotal technology for bridging Multimodal Large Language Models (MLLMs) with real-world AI hardware, such as AR glasses and embodied robots. However, current benchmarks are largely confined to a simplistic “immediate-query, immediate-answer” scenario. We argue that real-world interactions demand a more advanced capability for **Asynchronous Response**, where a model's output is delayed until specific visual conditions are met within a maximum temporal window. This paradigm fundamentally necessitates **fine-grained understanding**, as the model must vigilantly track the subtle, object-centric details that act as the visual triggers. This critical interplay between asynchronous response and fine-grained triggers has been largely overlooked.Therefore, we propose ASVBench, comprising 1452 videos and 4350 challenging Q&A pairs, as the first benchmark for evaluating Asynchronous Response to Fine-grained Triggers in streaming video.Specifically, we build a dual-axis evaluation framework for achieving this goal:(1) The temporal axis assesses models across three core scenarios: Immediate Q&A (responding to the present visual state), Prospective Q&A (responding when future visual triggers are met), and Continuous Q&A (continuously monitoring if triggers are met).(2) The cognitive axis requires the model to respond to visual triggers at different granularities, including recognition-level, comprehension-level, and reasoning-level triggers.ASVBench evaluation results reveal a significant gap between existing MLLMs and human performance, with accuracy plummeting as asynchronous temporal complexity increases. Detailed diagnostics indicate that current models suffer from micro-feature over-smoothing, memory eviction of latent goals, and impatience-driven premature responses, with early errors cascading in multi-shot streams. These insights expose critical bottlenecks in asynchronous responses to fine-grained triggers, providing a roadmap for next-generation streaming MLLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.