Mitigating Memorization In Text-To-Image Diffusion Models With Representation-Space Supervision
Abstract
Diffusion models are prone to memorization, i.e., reproducing individual training examples, which raises concerns about data privacy and copyright. Existing mitigation approaches primarily intervene through data preprocessing or specialized sampling procedures, leaving the role of the training objective comparatively underexplored. In this work we study whether changing the space in which diffusion predictions are supervised can mitigate memorization. We posit that direct data-space supervision contributes to memorization by encouraging model predictions to recover the exact training samples. Motivated by this view, we introduce Representation-Space Supervision (RSS), a general training framework that instead of supervising the models directly in data space, matches the model predictions to their targets in the feature spaces of pretrained visual representation models, which retain high-level perceptual and semantic information while reducing specific instance-level details in the training data that are essential for memorization. We evaluate RSS in controlled text-to-image supervised fine-tuning experiments, showing that it consistently and substantially improves the trade-off between generative performance and training-example reproduction relative to standard diffusion training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.