From Probes to Decisions: Multimodal Evidence Learning for Trustworthy Audio Reasoning
Abstract
Audio-language models can understand sound events and context in complex recordings, but being able to describe a recording does not imply having the evidence needed to answer a specific question. Question-guided acoustic measurements can provide targeted observations; however, even accurate observations may be insufficient to support the judgment at hand. We propose Multimodal Evidence Learning (MEL), a framework for trustworthy audio reasoning that connects question-relevant acoustic evidence acquisition with task-oriented evidence use. MEL decomposes a question's decision requirements into textual probes, guides acoustic tools to obtain observations with source context and scope, and combines these observations with the original audio and question for task decisions. The framework emphasizes what observations support and where their inferential limits lie, allowing local evidence to complement holistic audio understanding rather than be equated with task conclusions. We incorporate supervised learning into this process to strengthen the connection between a question's information needs and evidence acquisition, allowing local acoustic observations and holistic audio understanding to jointly inform task decisions. Experiments on MMAR and MMAU with two audio-language backbones of different sizes show that MEL improves overall accuracy over direct answering. Ablations and per-question analyses further reveal the contributions of evidence acquisition and learning, as well as how answer corrections and regressions jointly determine the net gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.