acceptodds
Under review as a conference paper at ICLR 2027

Strong Yet Learnable: Rethinking What Reasoning Distillation Needs

Abstract

Reasoning distillation transfers the reasoning capabilities of large language models to smaller students by supervising their intermediate decisions. However, stronger teachers do not necessarily produce better students, as increasingly sophisticated reasoning can become difficult for weaker students to absorb. On-policy self-distillation reduces this mismatch by using the same student as a privileged teacher with access to the verified outcome, referred to as hindsight. Yet its decisions rely on information unavailable to the deployed student and may therefore be difficult to reproduce causally. This raises a more fundamental question: what makes supervision worth learning? Useful supervision must provide knowledge beyond the student's current behavior while remaining learnable for the student. Crucially, these two requirements need not be satisfied by the same source. Hindsight provides a way to assess learnability: it gives the same student additional task information while keeping the model and reasoning state fixed, revealing whether a teacher signal remains compatible with the student under richer information. We therefore introduce HinD, an on-policy distillation framework that decouples knowledge provision from learnability assessment: a strong external teacher supplies new reasoning, while the student's hindsight response assesses its learnability and weights the corresponding supervision accordingly. Across multiple student scales and model families, HinD consistently improves mathematical reasoning and transfers to coding and knowledge-intensive reasoning, with gains that persist under larger sampling budgets and grow with teacher strength.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.