Making Audio Matter: MMALR for Benchmarking Audio Listening, Understanding, and Reasoning with Counter-Prior Construction
Abstract
Large audio-language models (LALMs) have made substantial progress across a wide range of audio tasks. To evaluate them consistently, Audio Question Answering (AQA) has become a mainstream paradigm with a unified cross-task interface. However, existing evaluations usually rest on an implicit assumption: a model's response is taken as its analysis of the audio. But in practice, the textual content alone may induce a preference for a particular answer, which arises from the model's textual prior. The effect of the prior becomes evident when audio is replaced with random noise: a considerable portion of questions can still be correctly answered, suggesting that high accuracy on existing benchmarks may not stem from genuine audio processing capabilities. Meanwhile, whether textual priors still affect model predictions with normal audio remains unexamined. To fill this gap, we introduce MMALR, a multi-domain, multi-task benchmark for audio listening, understanding, and reasoning that uses counter-prior construction to make audio evidence play a greater role in AQA. Specifically, MMALR anchors each correct answer in audio evidence and turns plausible textual predictions into distractors, making audio indispensable for reliable answering. We systematically evaluate open- and closed-source LALMs as well as human participants. Results show that under uninformative-audio conditions, MMALR yields the lowest model accuracy among all compared benchmarks. Furthermore, paired predictions under uninformative- and normal-audio conditions on MMALR show that models differ in their textual tendencies, while most models still exhibit consistent incorrect tendencies under the normal-audio condition. MMALR thus provides a more reliable benchmark for evaluating LALMs' audio listening, understanding, and reasoning, and a basis for analyzing their text-audio modality dependency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.