ConvoAudit: A Controlled Diagnostic Benchmark for Conversational Audio Understanding
Abstract
Understanding conversational audio is essential for intelligent systems to support effective communication and assistance. However, evaluating this capability requires assessing not only the content they recover but also how they track participants, interpret interactions, and locate events in time. Although existing benchmarks offer increasingly detailed assessments, aggregate scores can obscure selective weaknesses, while multiple factors always vary together in real-world conversational audio, making their individual effects difficult to isolate. In this work, we introduce ConvoAudit, a controlled diagnostic benchmark for conversational audio understanding, to address these challenges. Specifically, ConvoAudit evaluates dialogue content, interaction structure, and acoustic-temporal organization as three complementary dimensions to profile system capabilities under different conversational conditions. These conditions are realized via synthesized variants of shared source dialogues, thus enabling us to examine how capability profiles and performance gaps between transcription-based and direct-audio systems vary with conversational organization. Our evaluation on over fifty system configurations reveals that transcription-based and direct-audio systems can achieve similar scores on dialogue content yet differ substantially on interaction structure and acoustic–temporal organization, with the relative advantage of direct-audio systems varying across conversational conditions. Matched transcript controls further demonstrate that the usefulness of structural evidence depends on both its availability and its correctness. Fine-tuning on ConvoAudit's training dataset also enables open-weight models to outperform strong proprietary baselines, highlighting its value for targeted model improvement. Together, these analyses connect where systems struggle with controlled tests of which evidence improves their judgments, thereby supporting more informative evaluation of conversational audio understanding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.