Omni-InteractScope: Benchmarking Proactive and Agentic Interaction in Continuous Omnimodal Streams
Abstract
Omnimodal large language models (Omni-LLMs) are increasingly expected to assist users while observing continuous audiovisual streams. However, existing evaluations provide limited insight into how reliably these models connect stream evidence and user needs to timely, sustained interventions. We introduce __Omni-InteractScope__, a benchmark of 1,500 human-annotated, cross-reviewed interactions with 1,814 temporal targets across eight task groups. It links input modalities, event triggers, intervention types, and user needs, covering event-timed tool invocation and query-free assistance alongside notifications, warnings, state tracking, and continuous responses. Our experimental results on the proposed benchmark reveal three critical limitations of current Omni-LLMs: (1) __Timely tool invocation remains unreliable__: all 12 evaluated models produce an on-time first response for fewer than 40% of tool-invocation targets, and joint timing-and-tool accuracy reaches at most 31.1%. (2) __Query-free reminders remain poorly covered__: all evaluated models achieve lower timely reference-target recall on query-free reminders than on explicit questions, measured on their respective target sets. Reminder recall never exceeds 23.0%. (3) __End-to-end guidance coverage remains limited__: nine models cover none of the 26 guidance sequences in full, and the remaining three cover only one to three sequences each. Targeted controls further show that tool-invocation scores alone do not establish event grounding and that reminder coverage depends on response history and frequency. These findings motivate developing and evaluating streaming assistants for timely action, initiative under general user needs, and coverage of entire response sequences.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.