AV-RIFT: Task-Agnostic Transferable Adversarial Attacks on Audio-Visual Models
Abstract
Audio-visual models integrate auditory and visual cues to provide a comprehensive understanding of real-world scenes and have been applied to diverse tasks (e.g., event localization and question answering), largely benefiting from pretrained unimodal encoders that provide well-generalized representations. However, the open accessibility of these encoders also exposes a shared adversarial attack surface across downstream models, which remains unexplored by existing audio-visual attacks. This paper, for the first time, explores this threat by proposing AV-RIFT, a transferable audio-visual attack that solely utilizes publicly available encoders to mislead diverse downstream models. It jointly perturbs audio and video by disrupting semantic relationships both within and across modalities. Specifically, since heterogeneous audio-visual tasks commonly revolve around the semantics of sounding events shared by different instances, merely disrupting instance-level features on pretrained encoders does not necessarily alter the perceived event. AV-RIFT therefore introduces an intra-modal semantics-aware contrastive attack that steers each modality's representations away from semantically similar references and toward dissimilar ones. Moreover, to exploit the semantic consistency between audio and visual streams overlooked by independent unimodal perturbations, it further introduces an inter-modal misalignment attack that reduces the similarity between synchronized audio and visual representations in a projection space learned to bridge independently pretrained encoders. By exploiting the event-centric semantics and temporal synchronization inherent in audio-visual data, AV-RIFT is tailored to audio-visual models rather than being a direct adaptation of attacks designed for other modalities. Experiments on audio-visual tasks and models show that AV-RIFT consistently achieves state-of-the-art attack transferability and remains effective under existing defenses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.