MuLA-Bench: A Multilingual Long-Form Audio Understanding Benchmark via Multi-Tier Auditing
Abstract
Long-form audio performance is often summarized by context length and aggregate accuracy, obscuring how language, evidence, and task jointly shape difficulty. We introduce MuLA-Bench, a challenging benchmark of 5,038 open-ended questions over 1,732 in-the-wild recordings totaling 1,341 hours, covering 16 languages and eight domains. A balanced LanguageDomain semantic track supports controlled comparisons, while a complementary acoustic track preserves naturally occurring non-speech evidence. Evidence-grounded generation, shortcut checks, and language-expert review provide auditable questions without translating a shared source set or injecting target sounds. We evaluate eleven audio-language models and conduct targeted diagnostics of language, evidence, and temporal performance. Language rankings change across domains and tasks; acoustic-semantic performance gaps vary with the requested operation; and later evidence is associated with lower temporal accuracy but nearly flat long-range accuracy in a four-hosted-model cohort. Long-range retrieval is comparatively strong, while precise clock alignment and factual grounding of natural acoustic events remain fragile. MuLA-Bench thus exposes conditional failure patterns that a single long-context score does not capture. Anonymous code and benchmark artifacts are included in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.