acceptodds
Under review as a conference paper at ICLR 2027

Probing Concept Erasure in Diffusion Models via Initial Latent Optimization

Abstract

Concept erasure aims to suppress designated content in text-to-image diffusion models while preserving generation of other content. An illusion of forgetting occurs when suppression under ordinary prompts coexists with recoverable concept generation under adversarial inputs. Across the evaluated models, larger distributional discrepancy between erased and vanilla models' noise predictions during denoising is associated with lower target-concept detection rates. Motivated by this observation, we propose Initial Latent Variable Optimization (IVO) to assess concept recoverability through white-box attacks. IVO optimizes initial latents while keeping the erased model's parameters and original prompt fixed. Target-concept reference images provide the default initialization through DDIM inversion. A distribution matching loss and a direction calibration loss guide optimization using surrogate noise predictions. Experiments cover eleven concept erasure methods and three scenarios involving nudity, objects, and artistic styles. Average attack success rates reached 79.6%, 87.7%, and 91.8% for nudity, parachute, and Van Gogh, respectively. Both average attack success rates and CLIP scores exceeded those of the evaluated prompt-based attacks. A blinded human evaluation of prompt alignment also favored IVO over P4D and UDiff across these scenarios. The findings motivate including latent optimization in concept erasure evaluations and assessing recovery together with alignment to the original prompt.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.