State-Aware Contrastive On-Policy Distillation
Abstract
On-policy distillation trains a student on its own generations, while a single-teacher objective specifies only the distribution to imitate and no explicit extrapolation direction. We propose Contrastive On-Policy Distillation (COPD), which combines positive-teacher attraction and negative-teacher repulsion in a conditional KL objective on student-generated prefixes. We derive its closed-form virtual teacher and characterize the induced targets as an exponential-family path. The teacher pair supplies both an extrapolation direction and a coordinate system for targets and students. We localize the student along this path with a one-dimensional convex reverse-KL projection, whose residual measures off-axis mismatch. We extend COPD with State-Aware Contrastive On-Policy Distillation (SACOPD), which places each rollout's target a fixed coordinate increment ahead of the student's projection. With a Qwen3-4B student and 300 training rollouts, static COPD raises mathematics sample accuracy from matched OPD's to , compared with the positive teacher's . SACOPD reaches and improves both accuracy and coverage over COPD on all three code benchmarks. Together, these results connect effective teacher amplification with interpretable student localization and state-aware target selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.