LAMAR-Bench: Diagnosing Long-Form and Multi-Audio Understanding and Reasoning
Abstract
Audio-language models are increasingly able to process long recordings and multiple related audios. Existing benchmarks capture important parts of these settings, but the evaluation landscape remains fragmented, with fine-grained paralinguistic and speech-artifact analysis, long-range semantic dependencies, detailed audio-version comparison, and goal-conditioned decisions still underexplored. We introduce LAMAR-Bench, a capability-based benchmark that organizes long-form and multi-audio evaluation into three complementary families: fine-grained understanding, cross-audio relation reasoning, and decisions under goals and constraints. It contains 1,597 questions across 10 tasks, constructed from shared audio facts with traceable supervision, with up to 19 audio clips and 59.9 minutes of input per question. Evaluating 14 audio-language and omni-modal systems reveals substantial room for improvement: the best paralinguistic-analysis and acoustic-localization accuracies reach only 51.09% and 54.38%, respectively, while detailed cross-audio difference analysis remains challenging. We further conduct paired analyses over answer format, presentation order, and requested goals, revealing limitations in direct localization, order robustness, and goal adaptation. These findings highlight persistent gaps in evidence-grounded understanding and reliable goal-conditioned reasoning over long and multiple audio inputs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.