acceptodds
Under review as a conference paper at ICLR 2027

AVID: Benchmarking Audio-Visual Inconsistency Understanding for Omni-Modal Language Models

Abstract

Omni-modal large language models demonstrate strong content understanding in tasks such as captioning and question answering, yet their ability to understand fine-grained audio-visual inconsistencies remains largely unevaluated. To bridge this gap, we present AVID, a benchmark for audio-visual inconsistency understanding comprising 11.2K long-form videos with 39.4K annotated inconsistency events and 78.7K derived segment clips, with source-video-disjoint training and test splits. It supports detection, temporal grounding, classification, and reasoning across eight fine-grained inconsistency categories, including temporal shifts, semantic contradictions, and environmental conflicts. Compared with prior benchmarks focused on aligned events, forgery detection, or controlled multiple-choice diagnostics, it evaluates the localization and explanation of multiple inconsistencies in long videos, spanning low-level physical mismatches and high-level semantic and contextual conflicts. The benchmark is constructed through a scalable pipeline that combines content-type-aware temporal segmentation with an agentic framework for category selection and diverse audio-visual conflict injection. Modality ablations and processed-consistent controls show that strong performance requires cross-modal comparison and cannot be explained by manipulation artifacts. Extensive experiments and analyses of open- and closed-source models reveal persistent limitations in fine-grained inconsistency understanding. Using the training split, we develop AVID-Qwen, which improves over its base model across detection, temporal grounding, classification, and reasoning. Adding examples from the training split to general training data improves inconsistency detection while retaining general omni-modal capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.