acceptodds
Under review as a conference paper at ICLR 2027

Contrastive Representation Learning for Robust Concept Erasure in Diffusion Models

Abstract

Concept erasure edits pretrained text-to-image diffusion models to remove certain unwanted or unsafe concepts while preserving other capabilities. These erased concepts often stay suppressed under ordinary prompts, however they can be recoverable under adversarial inputs. We introduce CIRCE (Contrastive Image- Representation Concept Erasure), which reshapes the diffusion model’s internal representations through contrastive learning. CIRCE fine-tunes diffusion models to retain benign concepts in the representation space, as well as unlearn the harmful concepts by using them as a negative example. We evaluate CIRCE across objects, artistic styles, and nudity on Stable Diffusion and Qwen-Image-2.1. On UNLEARN-CANVAS, CIRCE achieves the highest mean across six erasure and retention metrics in our comparison. For nudity erasure, CIRCE remains robust across fixed jailbreak prompts, adaptive prompt search, and continuous embedding attacks, while pre- serving strong image quality on benign prompts. A human study with 24 annotators further finds the best observed preservation on general prompts for both backbones. Preservation on benign prompts close to the erased concept remains competitive with existing methods. Overall, our results show that directly reshaping internal representations provides a promising direction for robust concept erasure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.