Towards Distillation-Resistant Large Language Models: A Conditional Mutual Information Perspective
Abstract
Proprietary large language models (LLMs) embody substantial economic value and are generally exposed only as black-box APIs. However, adversaries can still exploit their outputs to extract knowledge via distillation. Previous studies have explored defense strategies, yet they primarily focus on text-based distillation, leaving the important logit-based counterpart largely unexplored. In this work, we analyze this problem from an information-theoretic perspective and present an effective solution. Specifically, we characterize distillation-relevant information in teacher outputs via the conditional mutual information (CMI) between teacher logits and input queries, conditioned on ground-truth labels. This quantity captures contextual signals beyond label supervision that can be exploited for knowledge extraction, motivating us to defend against unauthorized distillation via CMI minimization. To this end, we propose a learnable transformation matrix that purifies the original outputs and provide a theoretical analysis that establishes its effectiveness. We further derive a CMI-inspired anti-distillation objective to optimize this transformation, which effectively reduces distillation-relevant information while preserving output utility. Extensive experiments across multiple LLMs and strong distillation algorithms demonstrate that the proposed method significantly degrades distillation performance while preserving task accuracy, thereby effectively protecting models' intellectual property.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.