acceptodds
Under review as a conference paper at ICLR 2027

Learning What to Teach: Adaptive Hint Generation for On-Policy Self-Distillation

Abstract

On-policy self-distillation (OPSD) turns a language model into its own teacher by supplying it with privileged context (a hint) and distilling the resulting token-level guidance back into the unconditioned student. Existing methods often supply task-level hints that capture what a task generally requires, rather than what the current student gets wrong. We argue that the hint should instead be learned. We introduce EvoHint, which trains a dedicated hint generator to produce reflective hints: guidance conditioned on the student’s own attempt and optimized by whether it helps the student succeed. The generator draws on a persistent memory of effective hints, adapting retrieved hints to the observed attempt. Same-task student rollouts provide outcome feedback for generator training and select successful hints for memory updates, allowing guidance to adapt as the student learns. Experiments on agentic tasks (Search-QA and WebShop) and six mathematical reasoning benchmarks show consistent improvements over a fixed-hint self-distillation baseline across student model scales, with gains of up to 9.8 percentage points on agentic tasks and 1.8–3.1 points in average mathematical reasoning accuracy. Ablations further show that EvoHint’s gains hold across student objectives, that learned hints outperform frontier-model hints by up to 5.4 points, and that reusing generated hints as memory performs on par with frontier-model-built memory without a separate memory writer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.