Masked Self-Distillation: Internalizing the Chain-of-Thought in Language Models
Abstract
Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate traces dominate latency, memory usage, and serving cost, even though the final answer correctness is not causally related to the trace correctness and the trace length is not a reliable indicator of the problem complexity. This raises an obvious question: can the computation expressed in these intermediate tokens be internalized into the parameters of a language model, enabling it to produce answers with much shorter intermediate traces? We propose masked self-distillation, a knowledge-distillation based post-training framework in which copies of the same model are instantiated as teacher and student, and the student model is trained to internalize all or part of the intermediate trace, thus becoming more efficient at inference. We vary the fraction of the intermediate trace the student is trained to internalize, interpolating between full internalization (the student emits only the solution) and no internalization (the student emits the full trace before the answer). We conduct controlled experiments on two reasoning domains: math and graph coloring. We use the masked self-distillation framework to post-train Qwen3-4B and Qwen3-8B models. Our results demonstrate that this method can be used to improve task performance while increasing inference efficiency across various domains and model sizes. We systematically analyze whether the improved efficiency gains in the post-trained models generalize to out-of-distribution (OOD) problems. We find that the masked self-distillation models generalize well for in-domain OOD problems, and the masked self-distillation training does not induce catastrophic forgetting in the student model on out-of-domain problems. Furthermore, our ablation study shows that supervised fine-tuning can train models to produce shorter traces, but at the cost of generalization, highlighting the importance of on-policy training in masked self-distillation. Lastly, we show that this framework can also be used effectively in a dual-model setting where a larger reasoning model teacher supervises a smaller non-reasoning student model, resulting in improved task performance and inference efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.