Tangent On-Policy Distillation: Local-Direction Supervision Beyond Outputs
Abstract
On-policy distillation (OPD) reduces the train–test mismatch by training on student- generated trajectories, but most existing methods supervise output distributions and do not directly constrain how intermediate representations evolve along those trajectories. This raises the question: can local changes in hidden representations provide an effective supervision signal beyond output matching? We introduce Tangent On-Policy Distillation (TOPD), which constructs teacher and student token-by-layer hidden-state grids on the same student-generated rollout and aligns first finite-difference directions along adjacent response tokens and aligned network layers. TOPD combines this local-direction supervision with pointwise state matching, using cosine distance and detached, capped loss-ratio balancing to calibrate their scales. Distilling JustRL-DeepSeek-1.5B into DeepSeek-R1-Distill-Qwen- 1.5B on DAPO-Math-17K, TOPD achieves 58.66 Macro Avg@16 across AIME24, AIME25, and AIMO, outperforming the strongest prior distillation baseline by 4.16 percentage points (7.6% relative). Ablations support this design: removing temporal or depth supervision lowers Macro Avg@16 by 1.70 and 1.00 points, respectively, while the proposed cosine objective with detached, capped balancing outperforms fixed weighting, magnitude matching, differentiable balancing, and uncapped balancing. These results show that local representation directions provide an effective complement to output-level supervision in on-policy distillation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.