acceptodds
Under review as a conference paper at ICLR 2027

Projected Distillation: On Learning to Solve Hard Problems from Close, Correct Targets

Abstract

On-policy learning is central to post-training: training on a model’s own generations limits distribution shift and helps preserve existing capabilities. On hard problems, however, on-policy RL provides no reward-based learning signal when all attempts fail. On-policy distillation provides next-token supervision from a teacher, but on hard problems, if a rollout goes wrong early, subsequent supervision is conditioned on prefixes that already contain the mistake. Correct behavior can instead come from a gold or teacher solution, but such off-policy targets can differ substantially from model’s generations and, as we show, may erode existing capabilities. We therefore treat on-policyness as a continuum and propose **projected distillation**, which minimally edits the student’s failed attempts into correct targets. The method uses contrastive updates to favor corrections over failed attempts and refreshes these corrections as the student evolves, enabling near-on-policy learning even when all student attempts fail. Theoretically, we prove that the student’s learning regret grows with the edit size: it scales linearly with the horizon under minimal correction, as in DAgger, and quadratically under full-answer replacement, as in behavior cloning. Empirically, across two students (`Qwen3-8B` and `OLMo-2-7B`) and five domains spanning math, code, science, and tool use, projected distillation gives the highest area under the specialization–retention frontier (best on 8 of 10 model–task pairs). Our ablations highlight the importance of online corrections, contrastive loss, and, on-policyness: moving targets farther off-policy steadily worsens the specialization–retention trade-off, with retention falling from 95% to 75% when minimal corrections are replaced by student-agnostic gold solutions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.