SR-OPSD: Self-Referenced On-Policy Self-Distillation
Abstract
On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on student-generated trajectories, complementing reinforcement learning with sparse outcome rewards. Its self-teacher, derived from the student's current or exponentially averaged parameters and conditioned on additional context, evolves alongside the student and its rollout context distribution. The benefit of modifying this moving target depends on how target–student probability mismatches translate into updates. We propose Self-Referenced On-Policy Self-Distillation (SR-OPSD), which constructs a normalized geometric target from the self-teacher and a frozen initial policy, then minimizes the forward R\'enyi divergence from this target to the student. The interpolation coefficient controls the self-teacher's contribution, while the R\'enyi order controls the power weighting of target-to-student probability ratios in the gradient. For fixed contexts and target components, we establish a conditional variational characterization and derive the exact token-logit gradient, revealing how anchoring and projection jointly shape the effective update target. Experiments across scientific reasoning, tool use, mathematical reasoning, and code generation demonstrate strong performance across multiple model families and scales. Ablations further show that reference anchoring can improve or degrade performance depending on the projection objective, supporting the joint design of target construction and projection geometry.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.