acceptodds
Under review as a conference paper at ICLR 2027

Where Does the Audio Jailbreak Live? A Controlled Frequency-Depth Audit of Audio-Language Models

Abstract

Recent attacks and defenses for audio-language models (ALMs) assume that the jailbreak signal in an adversarial waveform concentrates in particular mel bands. We test this assumption with a frequency-depth occlusion protocol: we remove band-limited components of an additive perturbation in the STFT domain, re-run the victim, and score each condition on judge-adjudicated attack success and on the layer-wise divergence of decoder audio-span representations. The protocol is paired with controls that any band-localization claim should pass: bandwidth-matched random bands, equal-Hz and equal-energy partitions, matched-energy scattered removal, and uniform rescaling. For AdvWave-P against Qwen2-Audio (520 AdvBench prompts, attack success 0.77), a standard 8-band mel analysis points to the upper bands, but this locus does not survive the controls. A random contiguous region of the same bandwidth breaks the attack as effectively, every band is about equally load-bearing under an equal-energy partition, and a band's share of perturbation energy alone reproduces the occlusion ranking (Spearman 0.95). The attack does depend on how energy is removed: masking one contiguous band lowers success to 0.05 to 0.29, whereas removing the same energy uniformly leaves it at 0.75 to 0.79, and no single band or pair carries the attack even when it is re-optimized on that support. Applied to this attack, the band selections of ALMGuard and GRM do no better than matched random selections. In depth, band importance and representational divergence are aligned from the decoder entrance onward and remain so after adjusting for band energy, but we find no behavior-specific depth locus. The frequency findings replicate on three further prompt corpora. For this attack, the jailbreak is a broadband property of the jointly optimized perturbation, and band-selective claims should be reported against energy- and bandwidth-matched controls.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.