SelfCAD: Protecting Your Efficient Reasoning Capabilities via Self-Cautious Insertion
Abstract
Large reasoning models (LRMs) are valued for their accuracy and transparency, but publishing their reasoning traces also enables unauthorized distillation, raising copyright and security risks. Existing defenses either detect misuse after the fact or obscure reasoning traces, making them impractical when transparency is required. We propose SelfCAD (Self-Cautious Anti-Distillation), a lightweight post-processing defense that preserves the original reasoning content while hindering the performance and reasoning efficiency of distilled models. Our analysis shows that self-cautious sentences strongly affect both accuracy and efficiency: inserting such sentences can lengthen chain-of-thought reasoning while reducing accuracy in distilled models. SelfCAD therefore manipulates these components after traces are generated, requiring minimal extra cost and no additional training. Experiments on Llama and Qwen show that SelfCAD consistently increases student inference cost and reduces distilled-model accuracy. On Qwen with GSM8K, it increases response length by and reduces accuracy by 23 percentage points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.