ATTEST: Evidence-Grounded Reinforcement Learning for Long-Audio Temporal Understanding
Abstract
Temporal grounding is fundamental to audio understanding, yet large audio language models (LALMs) remain weak at it on long audio, where hallucinations compound and annotated data is scarce. Existing benchmarks cover only short clips or offer coarse, synthetic annotations. We introduce EgoSCAPE, a manually verified benchmark created from Ego4D with 700 real audios of durations up to 5 minutes, annotated with dense captions, sound sources, audio timelines, and temporal grounding queries. We further propose ATTEST, a GRPO-based reinforcement learning framework that improves temporal grounding without changing the base model's architecture or token vocabulary. ATTEST combines (i) a counterfactual grounding reward penalizing predictions that survive removal of their acoustic evidence, (ii) an attention-density reward encouraging generation to attend to its own claimed time spans, and (iii) token-wise advantages giving segment-level credit for long dense captions. ATTEST outperforms prior temporal grounding methods on existing dense captioning and grounding benchmarks, with fewer hallucinations. Together, EgoSCAPE and ATTEST establish a rigorous evaluation–training recipe that brings temporally grounded audio understanding closer to practical, long-form deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.