acceptodds
Under review as a conference paper at ICLR 2027

Concept Erasure for Text-to-Image Diffusion Models Against Image Probes

Abstract

Concept erasure aims to prevent text-to-image (T2I) diffusion models from generating unwanted concepts. Although T2I models can naturally accept images as input, existing erasure methods have limited the probing of erased models to only the text modality, e.g., “Generate a picture of a naked woman”. This makes existing erasure methods insufficient against image probes. In this paper, we address this research gap by proposing Image Probe Erase (IPErase), the first concept erasure method against image probes. IPErase adopts a frozen teacher (U-Net or DiT) model to construct denoising trajectories that extrapolate beyond the neutral prediction, away from the concept. It then trains a student model to match these trajectories using concept images as conditioning inputs through a frozen image prompt adapter (IP-Adapter), selectively updating the shared denoising backbone. On two widely used target concepts (nudity and Van Gogh style), our IPErase outperforms existing (text-based) erasure methods by (relative) 95.7%, 50.0%, and 91.8% against three increasingly challenging forms of image probes: black-box input condition, gray-box latent initialization, and white-box adversarial optimization. Further analyses show that despite relying solely on image supervision, our IPErase achieves perfect performance comparable to text-based methods against diverse text probes. This strong contrast in erasure generalizability between text-based methods to image probes and image-based methods to text probes suggests that images, which contain richer semantics, position target concepts much better than texts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.