Do Not Train on My Data: Unlearnable Data via Contrastive Reverse Distillation
Abstract
We propose **CoReDistill** (_**Co**ntrastive **Re**verse **Distill**ation_), a novel and efficient approach for making image data *unlearnable* from unauthorized generative-model training. Unlike existing methods that iteratively *optimize* each image toward an unlearnability objective, we introduce a sampling-based approach via contrastive decoding that directly *generates* unlearnable samples in a one-shot decoding process. We provide theoretical justification that this sampling procedure produces unlearnable samples, and empirically show that it substantially reduces processing time per image while maintaining strong unlearnability compared with optimization-based baselines. We further observe that the resulting protection transfers across different generative-model variants, providing empirical evidence of cross-model unlearnability. These properties make CoReDistill practical and scalable for real-world data protection involving diverse generative models and large-scale datasets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.