DistillBench: Evaluating Unauthorized Distillation and Its Defenses
Abstract
Knowledge distillation is a standard technique for compressing and specializing models. However, the same mechanism can also enable unauthorized capability transfer from proprietary models to an adversary model. Industry reports indicate that such extraction increasingly occurs at scale, but provide few reproducible details. Existing studies use different assumptions about attackers, API access, resource budgets, and acceptable utility loss, making defense effectiveness difficult to compare across settings. We introduce DistillBench, a benchmark providing a unified protocol for the joint evaluation of adversarial distillation and defenses. DistillBench supports fixed and adaptive data acquisition and measures capability transfer alongside acquisition and training costs, teacher utility, teacher similarity, and detectability. We evaluate four distillation methods and six defenses across acquisition budgets, using two teachers and two students on mathematical reasoning. We find that no method dominates across settings, and evaluating a defense against a single distillation procedure can misrepresent its effectiveness. Most tested interventions reduce teacher utility more than student performance, with protection often weakening as acquisition budgets grow. We further find that defense effects can depend more on the student than the teacher. Fingerprints are not reliably detected after distillation, and paraphrasing weakens detection based on teacher similarity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.