SourceShield: Source-Branch Adversarial Disruption for Inversion-Free Image Editing
Abstract
Text-guided image editing enables controllable visual content generation, but also introduces potential risks of misuse. Adversarial disruptions can mitigate such risks by slightly perturbing the input image. However, existing disruption methods have been limited to defending against training-based or inversion-based image editing, leaving the more efficient, inversion-free paradigm unexplored. In this paper, we identify that existing methods struggle to generalize to this paradigm because they overlook the role of the editing trajectory. As a result, early-stage disruptions may be attenuated by later solver updates, and output-stage disruptions can only capture the final disruption effect. To address these limitations, we propose SourceShield, the first disruption method against inversion-free image editing that disrupts the editing trajectory in the source branch. In particular, SourceShield does not choose the target branch due to its target-prompt-specific nature and slow optimization. Specifically, SourceShield disrupts the sum of source-conditioned velocity terms across solver time steps, to maximize the velocity differences between source and target branches. Experimental results demonstrate that SourceShield largely improves existing disruptions under unseen prompts, with 41.5% higher FID and 115.1% lower ImageReward. Additional results show its superior transferability to unseen models, resilience to countermeasures, and imperceptibility of perturbations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.