acceptodds
Under review as a conference paper at ICLR 2027

The Acoustic Stalemate: Benchmarking Large Language Models via Audio-based CAPTCHAs

Abstract

Audio CAPTCHAs represent one of the most challenging real-world tests for multimodal AI systems, requiring robust auditory perception and logical inference under diverse acoustic degradation. However, existing benchmarks largely focus on uni-modal language or vision tasks, leaving a critical gap in systematic evaluation of audio-language models. Thus, we introduce Au-Bench, the first interactive, open-source benchmark specifically designed to systematically evaluate the auditory perception and reasoning capabilities of large language models through 18 diverse audio CAPTCHA tasks spanning Content-based and Rule-based paradigms. We propose a dual-gradient complexity framework comprising Audio-Temporal Reasoning Depth (ATRD) to quantify the cognitive, multi-step reasoning burden, and Acoustic Perceptual Load (APL) to formally measure physical signal degradation through SNR, reverberation index, and spectral masking density. Our fully interactive web-based testing platform systematically evaluates state-of-the-art native large audio-language models and cascaded ASR-LLM systems under both zero-shot and Chain-of-Thought settings. Comprehensive experiments reveal that current foundation models perform substantially below human baselines, demonstrating acute vulnerability under both physical acoustic degradation and multi-step temporal reasoning while displaying highly distinct failure topologies. Au-Bench aims to shift the research frontier from conversationally fluent assistants toward robust, perceptually grounded agents capable of navigating real-world web interactions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.