Concept Erasure for Text-to-Image Diffusion Models Against Image Probes
Abstract
Concept erasure aims to prevent text-to-image (T2I) diffusion models from generating unwanted concepts. Although T2I models can naturally accept images as input, existing erasure methods have limited the probing of erased models to only the text modality, e.g., “Generate a picture of a naked woman”. This makes existing erasure methods insufficient against image probes. In this paper, we address this research gap by proposing Image Probe Erase (IPErase), the first concept erasure method against image probes. IPErase adopts a frozen teacher (U-Net or DiT) model to construct denoising trajectories that extrapolate beyond the neutral prediction, away from the concept. It then trains a student model to match these trajectories using concept images as conditioning inputs through a frozen image prompt adapter (IP-Adapter), selectively updating the shared denoising backbone. On two widely used target concepts (nudity and Van Gogh style), our IPErase outperforms existing (text-based) erasure methods by (relative) 95.7%, 50.0%, and 91.8% against three increasingly challenging forms of image probes: black-box input condition, gray-box latent initialization, and white-box adversarial optimization. Further analyses show that despite relying solely on image supervision, our IPErase achieves perfect performance comparable to text-based methods against diverse text probes. This strong contrast in erasure generalizability between text-based methods to image probes and image-based methods to text probes suggests that images, which contain richer semantics, position target concepts much better than texts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.