acceptodds
Under review as a conference paper at ICLR 2027

How to catch distillation

Abstract

Knowledge distillation transfers capabilities between language models, but determining whether a student was trained from a particular teacher remains difficult. We study whether a fingerprint embedded in a teacher can survive this transfer and reveal the training relationship. Our method associates naturally occurring teacher features with fixed patterns of token-level log-probability shifts and embeds these patterns through LoRA fine-tuning, with a distribution-preservation objective that limits changes to the teacher's behavior. A decoder queries a suspect model for output log probabilities on feature-active and feature-inactive contexts, measures alignment with the fingerprint, and compares that alignment with randomly generated fingerprints. The method requires access to the teacher during fingerprint construction but does not require the suspect model's weights or hidden states. Across the evaluated teacher–student pairs, fingerprints remain detectable after both logit-based and text-based distillation, including transfer across model families. Embedding increases teacher perplexity by 2.79% in the Qwen2.5-32B experiment. Additional evaluations examine a larger pool of 100,000 random fingerprints, 27 negative-control models, and continued student fine-tuning on clean text. In the tested fine-tuning configuration, detection persists after 10 million additional training tokens. These results support feature-conditioned fingerprinting as a practical direction for tracing language model distillation when the teacher can be modified before its outputs are used for training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.