acceptodds
Under review as a conference paper at ICLR 2027

Do Not Train on My Data: Unlearnable Data via Contrastive Reverse Distillation

Abstract

We propose **CoReDistill** (_**Co**ntrastive **Re**verse **Distill**ation_), a novel and efficient approach for making image data *unlearnable* from unauthorized generative-model training. Unlike existing methods that iteratively *optimize* each image toward an unlearnability objective, we introduce a sampling-based approach via contrastive decoding that directly *generates* unlearnable samples in a one-shot decoding process. We provide theoretical justification that this sampling procedure produces unlearnable samples, and empirically show that it substantially reduces processing time per image while maintaining strong unlearnability compared with optimization-based baselines. We further observe that the resulting protection transfers across different generative-model variants, providing empirical evidence of cross-model unlearnability. These properties make CoReDistill practical and scalable for real-world data protection involving diverse generative models and large-scale datasets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.