acceptodds
Under review as a conference paper at ICLR 2027

Layer Agreement for Target Reformulation in On-Policy Distillation

Abstract

On-policy distillation (OPD) supervises a student on its own trajectories, but large teacher–student distribution gaps can destabilize direct alignment with the teacher's final predictions. This raises a target-construction question: what distribution should the student learn at its current state? We find that directional agreement between teacher and student probability revisions across layers is strongly associated with smaller final-distribution mismatch, providing evidence of local behavioral compatibility. Motivated by this observation, we propose Layer Agreement for Target Reformulation (LaTR), which uses layer evidence to construct an intermediate alignment target. At each student-visited prefix, target construction uses two signals. Teacher guidance uses the final teacher–student probability gap to specify each candidate token's correction direction. Layer evidence uses intermediate-to-final probability changes to assess both models' support for that direction. Combining these signals yields a bounded adjustment to teacher logits that reshapes both candidate-token selection and relative target probabilities. The student then aligns with a teacher-anchored distribution informed by both models' local behavior to mitigate alignment instability caused by large teacher–student distribution gaps. We evaluate three students across seven mathematical reasoning and code-generation benchmarks. LaTR achieves the best reported overall averages and code scores among compared methods, with gains over the strongest baseline of up to 3.81 percentage points in the seven-benchmark average and 4.70 percentage points on LiveCodeBench-v6.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.