Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
Abstract
Distillation via reasoning traces exposes frontier models to adversarial third parties, who can bypass their guardrails and misappropriate their capabilities. Existing antidistillation defenses perturb the teacher's sampling distribution to hinder downstream student learning, but these perturbations also degrade the trace quality even when task accuracy is preserved, often leaving visible artifacts such as syntactic incoherencies and repetition chains. Beyond eroding user trust, such traces are statistically distinguishable, allowing adversaries to filter or resample them and weaken the defense. An effective defense must therefore limit how visibly it degrades the teacher. To address this gap, we formulate antidistillation as a max-min game whose constraint set explicitly encodes defense detectability. Satisfying both axes— degrading the student while remaining undetectable—may require more clever perturbations, akin to sparse but targeted adversarial attacks in robustness literature. We apply the same analogy to reasoning steps, targeting thought anchors: sentences identified in prior interpretability work as strongly influencing a model's output. We propose Thought Anchor RemovalS (TARS), which identifies and removes thought anchors, whose content is semantically redundant, so that the reasoning trace remains legible to the user. Modeling these removals as burst-deletions, we then provide formal guarantees on its detectability. Finally, we empirically demonstrate the efficacy of TARS in consistently hindering downstream distillation across teacher model scales (DeepSeek R1 Distill Qwen 7B, Gemma 3 12B), reasoning benchmarks (MMLU Pro, MATH, DeepMath), and student architectures (Llama 3.2 1B, Llama 3.2 3B, Gemma 3 1B, Qwen 3 1.7B) while preserving teacher performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.