Layer Agreement for Target Reformulation in On-Policy Distillation
Abstract
On-policy distillation (OPD) supervises a student on its own trajectories, but large teacher–student distribution gaps can destabilize direct alignment with the teacher's final predictions. This raises a target-construction question: what distribution should the student learn at its current state? We find that directional agreement between teacher and student probability revisions across layers is strongly associated with smaller final-distribution mismatch, providing evidence of local behavioral compatibility. Motivated by this observation, we propose Layer Agreement for Target Reformulation (LaTR), which uses layer evidence to construct an intermediate alignment target. At each student-visited prefix, target construction uses two signals. Teacher guidance uses the final teacher–student probability gap to specify each candidate token's correction direction. Layer evidence uses intermediate-to-final probability changes to assess both models' support for that direction. Combining these signals yields a bounded adjustment to teacher logits that reshapes both candidate-token selection and relative target probabilities. The student then aligns with a teacher-anchored distribution informed by both models' local behavior to mitigate alignment instability caused by large teacher–student distribution gaps. We evaluate three students across seven mathematical reasoning and code-generation benchmarks. LaTR achieves the best reported overall averages and code scores among compared methods, with gains over the strongest baseline of up to 3.81 percentage points in the seven-benchmark average and 4.70 percentage points on LiveCodeBench-v6.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.