Release-Time Anti-Distillation for Reasoning Models
Abstract
Frontier reasoning models are routinely distilled by third parties who query the model, collect (prompt, reasoning trace, answer) triples and fine-tune their own students on them. We study whether a model owner can *release* a model whose outputs remain fully useful to ordinary users yet are poor - or actively harmful - as distillation data. Our key observation is that the two audiences consume different parts of a response: users consume the final answer, whereas distillers extract most of their value from the reasoning trace. We therefore propose **Answer-Preserving Trace Poisoning (APTP)**: a release-time transform that leaves the answer bit-identical and rewrites only the visible reasoning trace. We instantiate three mechanisms with distinct gradient-level effects - *information removal* (hidden / summary / hollow traces), *gradient shortcut* (an early natural-language "intuition" that leaks the answer, so the answer loss is satisfied without reasoning), and *gradient conflict* (a **decoy** trace whose working coherently supports a different answer while the stated answer stays correct) - and we release AntiDistill-Bench, a 16.7k-problem benchmark with paired clean/poisoned traces and an evaluation protocol that measures (i) utility preservation, (ii) imperceptibility to casual readers and LLM judges, and (iii) distillation-scaling curves of students trained on the released data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.