acceptodds
Under review as a conference paper at ICLR 2027

MuLA-Bench: A Multilingual Long-Form Audio Understanding Benchmark via Multi-Tier Auditing

Abstract

Long-form audio performance is often summarized by context length and aggregate accuracy, obscuring how language, evidence, and task jointly shape difficulty. We introduce MuLA-Bench, a challenging benchmark of 5,038 open-ended questions over 1,732 in-the-wild recordings totaling 1,341 hours, covering 16 languages and eight domains. A balanced LanguageDomain semantic track supports controlled comparisons, while a complementary acoustic track preserves naturally occurring non-speech evidence. Evidence-grounded generation, shortcut checks, and language-expert review provide auditable questions without translating a shared source set or injecting target sounds. We evaluate eleven audio-language models and conduct targeted diagnostics of language, evidence, and temporal performance. Language rankings change across domains and tasks; acoustic-semantic performance gaps vary with the requested operation; and later evidence is associated with lower temporal accuracy but nearly flat long-range accuracy in a four-hosted-model cohort. Long-range retrieval is comparatively strong, while precise clock alignment and factual grounding of natural acoustic events remain fragile. MuLA-Bench thus exposes conditional failure patterns that a single long-context score does not capture. Anonymous code and benchmark artifacts are included in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.