acceptodds
Under review as a conference paper at ICLR 2027

FG-GUARD: FINE-GRAINED PERTURBATION DEFENSE WITH FULL PATH AUDIO ATTACK ANALYSIS

Abstract

Audio-Language Models (ALMs) exhibit a safety inconsistency: a harmful request may be rejected when spoken neutrally yet elicit an unsafe response when the same words are delivered differently. Existing attacks establish this vulnerability empirically, but do not provide a unified account of how an audio manipulation affects the safety of the generated answer. We develop a local, first-order analysis of the complete pathway from acoustic transformation to safety-relevant output. Signal-level transformations first induce structured changes in the log-Mel representation. The ALM and output Jacobians then propagate these changes through the model to a refusal-related opening-token probability. The resulting first-order expression quantifies the signed local change in this continuous proxy. The acoustic transformation determines the time–frequency pattern of the input change, while its effect on the proxy depends on the specific model and input. Motivated by two consequences of the analysis, we propose Fine-Grained Perturbation Defense (FG-Guard). The additive log-Mel formulation motivates a perturbation-based defense in the same representation space, while the joint time–frequency structure of the derived changes motivates a two-dimensional mask over the perturbation. The mask support is estimated empirically from jailbreak and transcription gradients. On Qwen2-Audio-7B-Instruct, FG-Guard reduces AdvWave attack success from 84.61% to 41.03% and AdvBench-Audio from 1.92% to 0.00%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.