FRA-Attack: Frequency-Domain Regularized Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs
Abstract
Multimodal large language models (MLLMs) remain vulnerable to transfer-based targeted attacks, where perturbations optimized on open-source surrogate encoders can generalize to closed-source MLLMs. A key challenge for improving adversarial transferability is to effectively capture the intrinsic visual focus shared across different models, such that perturbations align with transferable semantic cues rather than surrogate-specific behaviors. However, existing methods suffer from spatial-domain feature redundancy and surrogate-specific gradient signals, thereby hindering cross-model transferability. In this paper, we propose FRA-Attack, which addresses both challenges from a unified frequency-domain perspective. For feature alignment, a high-pass DCT objective concentrates on the high-frequency band of patch features that carries MLLMs' intrinsic visual focus, suppressing redundant global structures. For gradient optimization, we introduce a model-agnostic low-pass gradient regularizer on the frequency domain that modulates the surrogate gradient using only the geometric frequency coordinate, removing surrogate-specific high-frequency artifacts while preserving transferable low-frequency directions. Together, the two components form a unified frequency-domain treatment of transferability. Extensive experiments on flagship MLLMs across vendors show that FRA-Attack achieves superior cross-model transferability, particularly with state-of-the-art performance on GPT-5.4, Claude-Opus-4.6 and Gemini-3-flash.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.