Beyond Answer Quality: Prediction and Execution under Privileged Distillation
Abstract
Recent on-policy self-distillation (OPSD) methods have shown promise for improving language-model reasoning. A common form of OPSD uses privileged distillation, in which the teacher receives additional information, such as the correct answer, that is unavailable to the student. Yet studies of reasoning models also report cases where privileged supervision lowers answer accuracy and reduces the benefits of longer reasoning. Despite these observations, theoretical understanding of why such failures occur remains limited, making it difficult to understand when privileged supervision helps or hurts and how it should be improved. To address this gap, we develop a theoretical framework that separates two effects of privileged supervision: how well the student answers after completing the required reasoning steps (prediction), and whether it carries those reasoning steps through to completion (execution). We show that privileged supervision can make the student better at predicting the final answer once the required reasoning has been completed, while making it less likely to complete that reasoning on its own. As a result, better final-answer prediction can coexist with worse overall performance. We further show that this separation can emerge during finite distillation because learning at later reasoning steps depends on whether the student actually reaches those steps. Finally, we show a complementary effect on the teacher side: when privileged teachers are more likely to answer early, later reasoning steps are increasingly shaped by the teachers that continue reasoning. This changes what the student learns at those later steps, so different forms of privileged distillation can lead to different reasoning behavior. Experiments with Qwen3 and natural reasoning tasks provide qualitative support for these mechanisms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.