Diffusion Unlearning via Exact Continuous Mixture Scores
Abstract
Machine unlearning aims to mitigate exposure to sensitive training data and undesirable semantic concepts without retraining foundation diffusion models from scratch. While there are concept unlearning methods relying only on inference-time steering, data unlearning has predominantly relied on parameter fine-tuning, an approach that is computationally intensive and requires synthetic retain sets to prevent generation performance collapse. Emerging training-free alternatives circumvent retraining by steering reverse sampling trajectories away from target data, but they choose their repulsion force heuristically and avoid explicit computation of posterior terms as intractable. Conversely, we show that the mathematically principled repulsion weight is the instantaneous continuous-time posterior odds of entering forbidden data under the variance-preserving stochastic differential equation. We establish that unconditional data unlearning and conditional concept erasure share a common mathematical formulation, unifying both within a single training-free operator. We show that our proposed method, namely, Exact Mixture Score (EMIS) constitutes the exact spatial score of an explicit normalized mixture density with an a-priori closed-form leakage law and replaces heuristic time-window cutoffs with an uncertainty-aware continuous temperature schedule. Across unconditional pixel models (CelebA-HQ 256) and conditional text-to-image synthesis (Stable Diffusion on UnlearnCanvas for both style- and object-unlearning), EMIS not only outperforms heuristic training-free diffusion guidance methods but also matches the unlearning efficacy of full parameter fine-tuning while preserving generative fidelity on retained concepts with zero weight updates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.