A Single-shot Black-box Structural Hijacking Attack on Diffusion Models
Abstract
Diffusion models (DMs) have revolutionized text-to-image (T2I) and image-to-image (I2I) syntheses. However, they are widely recognized to be vulnerable to various malicious attacks, such as prompt injection, prompt evasion, and unsafe content generation. While most existing attack schemes require certain control over or access to the target DMs, we propose a single-shot black-box structural hijacking attack that requires no access to and minimal knowledge of these DMs. The core idea is to inject an adversarial perturbation into a source image, which enforces the perturbed image to have close intermediate key and value feature maps to those of a target image, typically a Not-Safe-For-Work (NSFW) image, in the denoising process. Once the perturbed image is processed by a victim with a DM, the generated image becomes corrupted in the still life case or even NSFW in the living subject case. Our attack scheme successfully sabotages state-of-the-art DMs by causing an average increase of over across three mainstream perceptual distances, including an animation-generating DM by attacking only the first frame. Hence this threat is rather significant and systemic, which calls for immediate countermeasures. NSFW Warning: This paper includes images that may be disturbing or explicit in nature.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.