acceptodds
Under review as a conference paper at ICLR 2027

Sight or Sound: Do Multimodal Large Language Models Have Blind Faith in Vision for Audio-Visual Understanding?

Abstract

Audio-visual understanding requires Multimodal Large Language Models (MLLMs) to reason over what is seen and heard, yet aligned benchmarks rarely reveal which stream controls an answer when the evidence conflicts. We construct the Audio-Visual Incongruity Corpus (AVIC), a question-conditioned probe that preserves questions while pairing answer-disagreeing streams and constructing source-attribution options. Across four MLLMs, outputs favor visual-source answers even under audio-focused instructions; conflict-aware prompting increases abstention but leaves substantial visual preference. An option-free evaluation filtered by isolated-source alignment on an independent dataset reproduces the visual preference. On a competence-controlled two-model subset, probes find locally decodable audio information under cross-modal conflict and prompt-dependent decision transfer; paired patching restores conflict-induced answer margins through late decision states. Modality masking confirms that auditory evidence can guide answers when visual access is blocked. Finally, an exploratory attention-intervention study on AVHBench tests early audio boosting and late visual suppression, revealing task-dependent effects and larger aggregate improvements from suppressing high-attention rather than low-attention visual tokens. The results connect a gap between local auditory information and decision influence to attention targets for improving audio-visual reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.