acceptodds
Under review as a conference paper at ICLR 2027

AutoAnchor: Stable Diffusion Unlearning Using Cross-Attention as a Manifold Surrogate

Abstract

Diffusion unlearning is essential for mitigating the generation of harmful or copyrighted content in text-to-image models. Current diffusion unlearning techniques determine the model update direction by either using alternatives of the target concept as anchors or using empty prompts. The anchor-based method relies on manually and semantically chosen anchors that risk biased unlearning, while the anchor-free method inherently suffers from imprecise unlearning due to unconstrained latent updates. In this work, we theoretically formalize diffusion unlearning under the manifold hypothesis and prove that lacking a manifold-proximal anchor can induce substantial normal-space drift that degrades unlearning performance. Motivated by this analysis, we propose AutoAnchor, a two-stage framework that automatically synthesizes manifold-proximal anchors. However, direct geometric manifold optimization is computationally intractable. To address this challenge, AutoAnchor introduces a novel cross-attention consistency loss, which serves as a highly efficient surrogate for manifold proximity. Experimental results demonstrate that AutoAnchor effectively achieves precise and unbiased unlearning across various state-of-the-art baselines, significantly improving targeted concept removal (by up to 28.17% in CLIP score) and non-target utility (by up to 4.17% in CLIP score). Moreover, AutoAnchor can also be easily integrated into existing diffusion unlearning methods to enhance their unlearning performance (by 34.48% for concept removal and 32.63% for utility on average).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.