After the Second Bell: Can Audio-Language Models Reason with Temporally Grounded Events?
Abstract
Audio-language models (ALMs) increasingly answer questions about speech, environmental sounds, and music, but temporal operations such as ordering, counting, and localization are often evaluated as separate skills. We study a more specific dependency: whether a model can identify a particular event occurrence, use its temporal position as a reference, and execute a downstream computation conditioned on that reference. We introduce (Temporal Event-Reference and Relational Reasoning), a diagnostic benchmark spanning speech, environmental sound, and music. The benchmark contains 67,466 questions in Task\ 1 across four source datasets and 2,955 matched questions in Task\ 2 over short MusicNet clips. Task\ 2 fixes two temporal intervals and asks models to count within each interval, determine interval overlap, and compute their intersection and union without double counting. Ground truth is generated from structured annotations and retained using cross-model agreement, followed by expert verification. Across the evaluation suite, models show a consistent gap between coarse temporal judgments and exact temporal counting: binary questions are often answered more reliably than numerical counting, while Task\ 2 intersection and union counting remain difficult even when interval-overlap recognition is substantially higher. We further find that Task\ 1 category scores are strongly affected by answer type and therefore should not be interpreted as a difficulty hierarchy. exposes a narrower failure mode: temporally grounded information can be recognized locally without being reliably carried into a subsequent relational computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.