ProactiveAudioBench: Evaluating Condition-Driven Responses in Audio Streams
Abstract
Continuous audio assistance requires models to execute a listening instruction as audio arrives, deciding when a target occurrence warrants a notification and when to remain silent. We introduce **ProactiveAudioBench**, covering sound events, speech, and music through Detection, Tracking, Counting, and Matching. These tasks organize first, repeated, and Nth-occurrence notification rules alongside targets specified by acoustic references. The benchmark contains 1,999 listening requests and 1,851 reference response opportunities, with negative requests requiring silence. Our evaluation framework supports simulated and native streaming, representing each request as a sequence of expected and emitted responses. Joint recall, joint precision, and negative-request false alarms assess the timing and content of notifications together with omissions and unnecessary replies. Across 23 system configurations and four tasks, we find protocol-dependent differences in error timing: simulated-streaming errors cluster before or near reference checkpoints, whereas native-streaming errors show a longer tail of delayed responses.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.