The Perceptual Bandwidth Tax: Matched Evaluation of Emotion and Accent as Jailbreak Surfaces in End-to-End Audio-Language Models
Abstract
End-to-end audio-language models can use not only the words in a request but also how those words are spoken. The same request can therefore reach the model in more forms, and its safety policy must remain consistent across them. We call this added consistency burden the perceptual bandwidth tax. PerceptBench evaluates this problem with reproducible speech variations rather than target-model gradients or request-specific optimization. Holding the request text fixed, we compare Direct TTS, six standard signal edits, and six emotion prompts on two unsafe benchmarks and seven open models. The emotion prompts span distinct affective regions and were selected through acoustic-cue screening; a separate accent study uses reference-cloned English speech after a rule-based pronunciation route proved unreliable. Our main measure is jailbreak coverage, the number of distinct harmful requests reached by each speech surface, with attack success rate and audio quality reported alongside it. Emotion prompting reaches more items than editing on six of seven models in JailbreakBench. On Step-Audio 2, it reaches 60 of 100 SORRY-Bench items, including 35 missed by both Direct TTS and editing, versus 23 reached by editing. Emotion-prompted audio also has higher predicted quality than edited audio (mean NISQA MOS 4.10 versus 2.81), and most of its successes survive a simulated quality filter. Reproducible changes in how a request is spoken can thus expose safety failures missed by direct speech and conventional signal edits.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.