acceptodds
Under review as a conference paper at ICLR 2027

From Fooling CNNs to Fooling MLLMs: Generating Unrestricted Adversarial Examples via Diffusion-Based Style Injection

Abstract

White-box adversarial attacks that directly backpropagate on multimodal large language models (MLLMs) or API query-based black-box attacks against MLLMs are both resource-intensive. Motivated by the observation that MLLMs suffer from difficulties in recognizing images of non-common styles, we propose a transferable StyleUAE to explore the vulnerability of MLLMs to stylized adversarial examples. StyleUAE is a low-resource attack that generates unrestricted adversarial examples (UAEs) on small models. To improve the transferability of UAEs, we first propose a diffusion-based style injection method to weaken the recognition ability of MLLMs, and then jointly attack a CNN and the cross-attention maps of the diffusion model in latent space to further disrupt discriminative cues. Considering that excessively stylizing and disrupting latent space may generate deformed UAEs, we leverage the harmonized image prior to reduce noticeable stylization, and further design structure-preserving components to preserve the original structure. Extensive experiments show that, compared with state-of-the-arts, StyleUAE achieves average ASRs (gains) of 77.5% (+45.8%), 80.9% (+56.7%), 77.2% (+21.4%), and 81.3% (+9.8%) when attacking open-source MLLMs, commercial MLLMs, zero-shot vision-language models, and visual unimodal classifiers, respectively. In addition, our UAEs exhibit satisfactory appearances and achieve competitive image quality scores. The anonymous code repository is available at https://anonymous.4open.science/r/StyleUAE.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.