Benchmarking Audio-Language Models on Real-World Long-Form Audio Understanding
Abstract
Large audio language models (LALMs) have made remarkable progress at the second- and minute-level scale, yet understanding hour-level audio remains a fundamental bottleneck. Existing benchmarks rely on short clips or artificially concatenated segments, and so fail to assess long-range comprehension in real-world scenarios such as podcasts or lengthy broadcasts. We introduce VoiceGiraffe, a bilingual, open-domain benchmark for extreme long-context audio understanding. It comprises 1,500 curated triplets under a dual-level taxonomy: single-hop questions for fine-grained grounding and acoustic perception, and multi-hop questions requiring cross-segment evidence aggregation. We evaluate a broad suite of open-source and proprietary LALMs against human performance, and complement the primary multiple-choice protocol with open-ended evaluation that removes candidate options. Three findings emerge. First, VoiceGiraffe is far from saturation, and only one end-to-end LALM surpasses the human reference. Second, inference paradigms are model-dependent. End-to-end inference favours models with native long-context capacity, and cascaded aggregation stabilizes models overwhelmed by hour-scale audio. Third, long-range memory persistence is the key bottleneck, where current models reason over prominent evidence once localized but fail to retain event states over hour-scale context. VoiceGiraffe thus offers a challenging, diagnostic testbed for long-form audio understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.