Projected Distillation: On Learning to Solve Hard Problems from Close, Correct Targets
Abstract
On-policy learning is central to post-training: training on a model’s own generations limits distribution shift and helps preserve existing capabilities. On hard problems, however, on-policy RL provides no reward-based learning signal when all attempts fail. On-policy distillation provides next-token supervision from a teacher, but on hard problems, if a rollout goes wrong early, subsequent supervision is conditioned on prefixes that already contain the mistake. Correct behavior can instead come from a gold or teacher solution, but such off-policy targets can differ substantially from model’s generations and, as we show, may erode existing capabilities. We therefore treat on-policyness as a continuum and propose **projected distillation**, which minimally edits the student’s failed attempts into correct targets. The method uses contrastive updates to favor corrections over failed attempts and refreshes these corrections as the student evolves, enabling near-on-policy learning even when all student attempts fail. Theoretically, we prove that the student’s learning regret grows with the edit size: it scales linearly with the horizon under minimal correction, as in DAgger, and quadratically under full-answer replacement, as in behavior cloning. Empirically, across two students (`Qwen3-8B` and `OLMo-2-7B`) and five domains spanning math, code, science, and tool use, projected distillation gives the highest area under the specialization–retention frontier (best on 8 of 10 model–task pairs). Our ablations highlight the importance of online corrections, contrastive loss, and, on-policyness: moving targets farther off-policy steadily worsens the specialization–retention trade-off, with retention falling from 95% to 75% when minimal corrections are replaced by student-agnostic gold solutions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.