It Takes Two to Tango: Backdoor Attack for Multimodal Diffusion Models via Cross-modal Collaborative Triggers
Abstract
With the increasing popularity of decoupled cross-attention-based Text-to-Image (T2I) Diffusion Models (DMs) in image editing tasks, maintaining their integrity and reliability is becoming an important issue. However, when dealing with multimodal feature fusion, such DMs rely heavily on two decoupled cross-attentions, which inevitably leak feature-projection information, making them vulnerable to exploitation by malicious adversaries. In light of this vulnerability, this paper introduces a novel Cross-Modal Backdoor Attack (XMBA) targeting these decoupling-based (i.e., decoupled cross-attention-based) T2I DMs. Unlike backdoor attacks on T2I DMs, which rely solely on text triggers, XMBA focuses on scenarios that involve the joint use of multimodal triggers. To enhance the effectiveness of the attack, XMBA adopts a target inversion method, which inverts the target image into a target text vector through the reverse process of DMs, improving the alignment between target texts and images. Meanwhile, based on our projection alignment approach, XMBA modifies the parameters of the text and image projection matrices to align trigger features with those of their corresponding targets in the latent space, enabling backdoor injection without additional training data. Comprehensive experiments demonstrate that XMBA can effectively attack decoupling-based T2I DMs in just 1.6 seconds, achieving a success rate of up to 99.5% without compromising the model's generative ability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.