acceptodds
Under review as a conference paper at ICLR 2027

AUDIOTRACE: EVIDENCE-GROUNDED QUESTION ANSWERING OVER LONG AUDIO

Abstract

Evaluations of large audio-language models (LALMs) have expanded beyond minute-scale clips to include recordings lasting several hours. Yet current long-audio QA evaluations face two key limitations. (1) Inflated scores can mask weaknesses in fine-grained audio understanding and reasoning. (2) Answer correctness is often assessed without verifying whether models identify the precise temporal spans and factual content supporting their predictions. To address these limitations, we introduce AudioTrace, a benchmark that disentangles answer correctness from evidence grounding in long-audio QA. The dataset comprises 120 Chinese and English recordings lasting 30–90 minutes each (99.8 hours total), with 1,033 pairs of multiple-choice and open-ended questions across five question categories. Our construction pipeline combines automated annotation with human review, yielding event-level temporal and content evidence for each answerable question. The evaluation protocol has three levels with independently generated responses: L1 assesses answer correctness alone; L2 additionally requires temporal localization of supporting evidence; and L3 further requires semantic grounding of that evidence. Experiments reveal a substantial gap between answer correctness and evidence-supported correctness. Across 17 model–paradigm configurations, mean L1 answer accuracy is 84.8% for multiple-choice QA and 43.8% for open-ended QA, whereas mean L3 joint accuracy is 22.1% and 10.9%, respectively. These results show that a correct answer alone does not establish accurate understanding of the supporting audio or reliable reasoning from it. Evaluating trustworthy long-audio QA therefore requires finer-grained checks of the temporal location and factual content of supporting evidence. The code and dataset are publicly available at https://anonymous.4open.science/r/open_audiotrace.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.