Is Preventing Model Distillation Possible? A Theoretical Perspective
Abstract
Proprietary AI models are publicly accessible via an API. Third parties can use this limited black-box access to help develop their own AI models by extracting training data. This practice is known as distillation and measures to prevent it are known as anti-distillation. We study this problem from a theoretical perspective and ask is it possible to provably prevent distillation of the entire model? We formalize this problem in the language of learning theory. We show that the feasibility of anti-distillation depends critically on the computational hardness of the underlying agnostic PAC learning problem and on the statefulness of the server. For some model classes and query distributions, provable defenses exist; for others, undetectable attacks exist.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.