Your T2I model's embeddings know about memorization and enable broad unlearning
Abstract
Text-to-image (T2I) diffusion models are known to memorize training examples, reproducing near-verbatim copies at inference time and raising privacy and copyright concerns. This troubling phenomenon has given rise to a plethora of methods to mitigate memorization, either via prevention measures, specialized inference procedures, or post-hoc unlearning methods. In this work, we make a surprising fundamental observation about memorized examples in text-to-image diffusion models: a simple logistic regression classifier on top of embeddings can predict with very high accuracy which examples are memorized, suggesting the existence of shared representational structure among memorized examples, which is distinct from non-memorized examples. We also show evidence that the relationship between memorization and the structures we identify in representation space is causal. The existence of this representational structure, aside from being scientifically interesting in its own right, is practically important as it opens the door to a new problem setting we refer to as “broad unlearning” of all memorized examples from knowledge of only a subset of them; a realistic setting that prior work has overlooked. Motivated by on our findings, we propose a new method called BLUR that is significantly more effective than prior work at reducing the memorization of both known and unknown (to the unlearning algorithm) memorized examples, while preserving the model's ability to generate high quality and diverse images.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.