acceptodds
Under review as a conference paper at ICLR 2027

Benchmarking Audio-Language Models on Real-World Long-Form Audio Understanding

Abstract

Large audio language models (LALMs) have made remarkable progress at the second- and minute-level scale, yet understanding hour-level audio remains a fundamental bottleneck. Existing benchmarks rely on short clips or artificially concatenated segments, and so fail to assess long-range comprehension in real-world scenarios such as podcasts or lengthy broadcasts. We introduce VoiceGiraffe, a bilingual, open-domain benchmark for extreme long-context audio understanding. It comprises 1,500 curated triplets under a dual-level taxonomy: single-hop questions for fine-grained grounding and acoustic perception, and multi-hop questions requiring cross-segment evidence aggregation. We evaluate a broad suite of open-source and proprietary LALMs against human performance, and complement the primary multiple-choice protocol with open-ended evaluation that removes candidate options. Three findings emerge. First, VoiceGiraffe is far from saturation, and only one end-to-end LALM surpasses the human reference. Second, inference paradigms are model-dependent. End-to-end inference favours models with native long-context capacity, and cascaded aggregation stabilizes models overwhelmed by hour-scale audio. Third, long-range memory persistence is the key bottleneck, where current models reason over prominent evidence once localized but fail to retain event states over hour-scale context. VoiceGiraffe thus offers a challenging, diagnostic testbed for long-form audio understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.