Mini-OPD: On-Policy Relation Distillation
Abstract
On-policy distillation (OPD) offers a natural remedy for exposure bias: rather than imitating a teacher on static data, a student is supervised on states it visits itself. Yet making the states on-policy does not make the supervision richer. Existing OPD still observes the teacher mainly through its next-token distribution, leaving the context processing behind that prediction largely underdetermined—especially after the student's trajectory has gone wrong. We introduce Mini-OPD, which extends on-policy action matching to on-policy state-processing alignment. Alongside output reverse KL, it distills response-level Q–Q, K–K, and V–V self-relations from one fixed teacher–student layer pair on the student's realized response. These relation rows inhabit a shared token-position simplex, enabling heterogeneous models to be compared without learned feature bridges when positions are aligned. The objective admits either KL order and uses the teacher-weighted branch by default; both targets use the same student rollout. Across AIME 2024, AMC, OlympiadBench, LiveCodeBench, and SWE-Bench, Mini-OPD is the strongest student method, improving over matched output-only OPD by 1.1–4.0 points on every evaluation. In a representative run, it reaches the low-divergence regime 3.41× earlier and attains a 66% lower reverse KL. These results suggest that correcting where a student learns is only part of on-policy distillation; enriching what it learns at those states matters as well.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.