acceptodds
Under review as a conference paper at ICLR 2027

CabinBench: Benchmarking What Long-Audio Agents Actually Heard

Abstract

Audio understanding has advanced on clips of seconds, where the answer and its evidence arrive in the same model call. Long recordings break that pairing: the evidence for a question may last three seconds inside an hour, so a correct answer may have been heard or guessed, and benchmarks that hand the model the whole recording and score only the answer cannot say which. A vehicle cabin is a natural testbed: over a trip of tens of minutes, several occupants talk over road noise, phone calls and navigation prompts cut in, and plans change, so understanding the scene means finding, tracking, and attributing evidence. CabinBench is, to our knowledge, the first benchmark that scores the listening process itself rather than the answer alone: the agent decides where to listen under a budget, must cite audio for every claim, and is credited only for citations that the harness’s delivery log confirms. Its recordings are synthesized from fictional scripts by design, which makes the ground truth fully controllable: every question carries an exact evidence contract, and any sentence can be removed, replaced, or moved to produce a counterfactual recording whose truth is known. On 200 recordings and 2,608 questions, a 30B-parameter open omni-modal model answers 16 to 20% of questions yet establishes its evidence in under 1% of runs; most of its correct answers never heard the sentence that settles the question, and a tenth of its submissions cite audio it was never given. Controllable data, an agent in the loop, and process-level scoring turn listening from an assumption into a measurement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.