Can Open-Weight LLMs be Protected from Distillation?
Abstract
LLM providers are concerned about their model’s ability being learned by competitors via distillation, which improves a model by training on outputs from another stronger model. To stop distillation, prior work post-processes LLMs’ raw reasoning traces before returning them to users to protect commercial APIs. These techniques, however, cannot protect open-weight LLMs, whose reasoning traces, decoding process, and model weights are completely controlled by the distillation attacker. As open LLMs rapidly approach the frontier, protecting them from being distilled emerges as an important problem that is under-explored. Motivated by it, we propose an antidistillation recipe that fine-tunes LLMs to reason in an encoded vocabulary while answering in their original vocabulary. The key idea is to keep reasoning traces useful to the model itself, but almost impossible for other models to learn from. Our encoded Qwen3-8B and QwQ-32B produce non-readable reasoning, making it detrimental for distillation. Thus, the attacker can only distill on final readable answers of our models, which is shown to be ineffective. We also evaluate 7 adaptive attackers that try to recover original reasoning from our encoded reasoning to distill from. The tested recovery attempts either recover at most 11% of the encoded–readable token mapping, or directly drop non-trivial utility of our model, making it a poor distillation teacher. While mitigating distillation, our encoded models still maintain 88%–99% of the original model’s utility scores on GSM8K, MATH-500, MMLU-Pro math, and AIME 2025. As a first step towards antidistillation for open LLMs, this work aims to open up possibilities for providers to release strong models’ weights while limiting distillation by competitors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.