EvaOmni: An Evidence-Grounded Benchmark for Audio-Visual Reasoning
Abstract
Audio-visual reasoning requires models to locate key audio and visual evidence and integrate information across modalities. However, existing benchmarks that rely primarily on multiple-choice questions typically score models solely on answer correctness, making it difficult to assess whether they can identify the evidence supporting their answers. Moreover, textual or unimodal shortcuts may allow models to answer correctly without analyzing the necessary audio-visual evidence, further weakening these benchmarks’ ability to assess audio-visual reasoning. To address these issues, we introduce EvaOmni, an evidence-grounded benchmark comprising 420 videos and 1,200 question-answer pairs across 15 capabilities and 3 levels: retrieval and grounding, basic understanding, and complex reasoning. Experts annotate audio and visual evidence and construct video-grounded distractors; text-only, audio-only, and visual-only checks filter questions that can be answered without both modalities. Models must return both answers and supporting evidence intervals. We propose the Coverage Grounded Score (C.G.) to measure how much of the annotated evidence a model identifies, and the Bottleneck Grounded Score (B.G.) to penalize poor temporal grounding in either modality; both incorporate answer correctness. Across 7 open-weight and 15 proprietary Omni LLMs, the strongest model achieves 80.0% answer accuracy but only 35.8% B.G., compared with 94.3% for human experts. At Level 3, its 79.1% accuracy contrasts with 12.4% B.G. These results show why audio-visual reasoning should be evaluated through both answers and the supporting evidence. We will publicly release EvaOmni to support the continued evolution of evaluation systems and further advance model development.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.