acceptodds
Under review as a conference paper at ICLR 2027

MS-OmniBench: When Listening Across Video Streams Is Necessary

Abstract

A camera may record the sound of an impact while its cause is visible only from another viewpoint. However, existing benchmarks for Omni-modal Large Language Models (Omni-LLMs) offer limited evaluation of reasoning that requires both sound and complementary information across streams. To address this gap, we introduce MS-OmniBench, a benchmark designed around questions that jointly require acoustic evidence and multiple streams. MS-OmniBench comprises 3,109 human-reviewed or human-authored questions across diverse real-world settings, organized into twelve subtasks spanning four families: synchronous cross-stream understanding, asynchronous cross-segment reasoning, audio-anchored retrieval, and cross-modal consistency verification. Answer relevant evidence is distributed across camera viewpoints, source-specific audio tracks, and non-adjacent temporal segments, with annotated evidence chains identifying the supporting streams and time ranges. We evaluate seventeen open-weight and proprietary Omni-LLMs and four existing agentic systems. The best proprietary and open-weight end-to-end models achieve 61.9% and 52.0% accuracy, respectively, while the strongest existing agentic system reaches 49.8%. We further introduce MESA, a simple Multi-stream Evidence Search Agent that organizes observations into per-stream event timelines and uses them to guide targeted media inspection and cross-stream answer verification. With Gemini 3.8 Flash, MESA reaches 65.6% accuracy, improving by 3.7 percentage points over the same backbone in the end-to-end setting. MS-OmniBench provides a testbed for studying how models use sound to connect complementary observations across streams. The benchmark will be released at **MS-OmniBench**.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.