On-Policy Self Distillation As Automatisation: When and Why Thinking Collapses in Reasoning Models
Abstract
On-policy self-distillation (OPSD) is a promising post-training technique in which a model teaches itself by distilling from a copy which is granted , such as the ground-truth answer to the task. It outperforms verifier-based methods like GRPO while making models more reasoning token efficient. However, on some tasks OPSD cuts reasoning so aggressively that accuracy drops sharply, a phenomenon called . Using the framework of usable, or -, information, we find that the longer and more relevant to the answer the privileged context is, the more the student shortens its reasoning, yet this shortening turns into collapse only the model generated that context itself. To explain , we draw on a classic distinction from cognitive neuroscience between slow, deliberate thinking () and fast, practiced responses (). OPSD can make reasoning automatic: the model stops extracting usable information about the correct answer as it reasons, and terminates earlier. Such automatic reasoning cannot adapt its effort to the task, leading to performance degradation as task difficulty increases, even outside the training domain. We additionally uncover a previously undocumented failure mode of OPSD: students inherit the teacher's trust in context, which makes them easier to misalign when that context is malicious, a risk for robustness and safety. Thinking collapse is thus practice-driven automatisation rather than a training fault.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.