acceptodds
Under review as a conference paper at ICLR 2027

TILDE: TILt-based Distributional Erasure for Concept Unlearning

Abstract

Concept unlearning in text-to-image diffusion and rectified-flow models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must reliably suppress unwanted concepts after training. Existing methods often remove the target concept effectively, but practical unlearning also requires an equally fundamental property: the unlearned model should retain quality, diversity, and semantic coverage on benign generation. The gold standard is a retain-only model trained from scratch without the unwanted data. However, common erasure objectives do not specify which post-unlearning distribution should approximate this reference, leaving retention as an implicit byproduct of heuristic updates. We propose TILDE, TILt-based Distributional Erasure, which formulates concept unlearning as a distributional alignment problem: the desired target is the minimum-deviation conditional distribution from the pretrained model under a forgetting constraint. This energy-tilted, anchor-free target downweights concept-expressing images while preserving benign relative mass for each prompt. We approximate this target with residual -GFlowNet training for diffusion and rectified-flow generators, learning only a score correction relative to the pretrained sampler via a modular, differentiable concept energy. Across objects, artistic styles, characters and NSFW content, TILDE achieves strong forgetting while improving retention and distributional fidelity over prior baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.