acceptodds
Under review as a conference paper at ICLR 2027

Tangent On-Policy Distillation: Local-Direction Supervision Beyond Outputs

Abstract

On-policy distillation (OPD) reduces the train–test mismatch by training on student- generated trajectories, but most existing methods supervise output distributions and do not directly constrain how intermediate representations evolve along those trajectories. This raises the question: can local changes in hidden representations provide an effective supervision signal beyond output matching? We introduce Tangent On-Policy Distillation (TOPD), which constructs teacher and student token-by-layer hidden-state grids on the same student-generated rollout and aligns first finite-difference directions along adjacent response tokens and aligned network layers. TOPD combines this local-direction supervision with pointwise state matching, using cosine distance and detached, capped loss-ratio balancing to calibrate their scales. Distilling JustRL-DeepSeek-1.5B into DeepSeek-R1-Distill-Qwen- 1.5B on DAPO-Math-17K, TOPD achieves 58.66 Macro Avg@16 across AIME24, AIME25, and AIMO, outperforming the strongest prior distillation baseline by 4.16 percentage points (7.6% relative). Ablations support this design: removing temporal or depth supervision lowers Macro Avg@16 by 1.70 and 1.00 points, respectively, while the proposed cosine objective with detached, capped balancing outperforms fixed weighting, magnitude matching, differentiable balancing, and uncapped balancing. These results show that local representation directions provide an effective complement to output-level supervision in on-policy distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.