MIRAGE: Raising the Cost of Malicious Image Editing via False Moderation
Abstract
The proliferation of AI-powered image editing systems raises serious concerns about personal images being arbitrarily manipulated at scale with minimal effort and a lower entry barrier. Prior works on image immunization require access to the model weights and the image editing prompt, which significantly limits their use, especially against powerful commercial black-box image editors such as GPT-Image, Gemini Flash Image (Nano Banana), and Grok Imagine. To address this, we take a system-level view of the problem and identify a previously unexplored protection surface common to most major commercial systems: pre-generation safety moderation. Rather than directly targeting the image-editing model, we target these moderation classifiers to flag immunized images as policy-violating, triggering an automatic refusal. We operationalize this by adding adversarial perturbations to align our image to policy-violating concepts in the representation space. Under the same level of perturbation, our approach, MIRAGE, significantly improves immunization success rate from less than 15% achieved by the best prior technique to more than 88% on multiple closed-source image editing APIs. Our approach is simple, prompt-agnostic, and effective, offering a practical path towards protecting personal images from unauthorized AI-powered editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.