acceptodds
Under review as a conference paper at ICLR 2027

Anchor-Conditioned On-Policy Distillation for Concept Erasure in Text-to-Image Generative Models

Abstract

Concept erasure in text-to-image (T2I) generative models aims to suppress specified concepts while preserving non-target generation capabilities. However, existing methods face a persistent trade-off between effective erasure and non-target preservation. Motivated by recent advances in on-policy distillation (OPD), we investigate whether supervision at student-visited states can improve this trade-off. We first apply OPD using existing concept-erased models as frozen teachers and find that it strengthens target suppression while largely preserving non-target generation capabilities across multiple erasure methods. Although effective, this procedure requires a separately trained concept-erased teacher. We propose Anchor-Conditioned On-Policy Distillation (AC-OPD), which instead uses a frozen copy of the original pretrained T2I model conditioned on a non-target anchor prompt. Specifically, the student generates trajectories under a target prompt specifying the concept to be erased and matches the frozen model’s anchor-conditioned transition means at student-visited states within an early generation window. Restricting supervision to this window yields a better erasure–retention trade-off than distillation over the full trajectory. Experiments on nudity and IP-character erasure demonstrate strong target suppression while largely preserving non-target text–image alignment and compositional generation capabilities. Compared with off-policy anchor-conditioned distillation, AC-OPD better preserves generation performance and reduces sensitivity to anchor choice. Code is available at https://anonymous.4open.science/r/AC-OPD-06F5.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.