acceptodds
Under review as a conference paper at ICLR 2027

Trajectory-Refined Distillation

Abstract

On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student’s own rollouts. In this work, we identify a common structural cause underlying OPD, which we call prefix failure. Under prefix failure, dense per-token supervision induces a bimodal teacher mixture and fragmented gradients that token-level loss truncation or reweighting fail to address. We thus propose Trajectory-Refined Distillation (TRD), a trajectory-level correction method that revises the student’s rollout under the teacher guidance while within on-policy support. By correcting problematic prefixes before distillation, TRD mitigates prefix failure at its source. Moreover, TRD improves exploration by exposing the student to alternative valid derivations under teacher guidance, even when the original rollouts are already correct. TRD can also be applied to on-policy self-distillation (OPSD), a parameter-sharing variant that uses the student model conditioned on privileged information as the teacher. TRD improves average math accuracy while remaining competitive on code. Under OPSD, it solves AMOBench questions left unsolved by the Qwen3-8B base model, nearly doubling the strongest dense-KL baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.