acceptodds
Under review as a conference paper at ICLR 2027

SR-OPSD: Self-Referenced On-Policy Self-Distillation

Abstract

On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on student-generated trajectories, complementing reinforcement learning with sparse outcome rewards. Its self-teacher, derived from the student's current or exponentially averaged parameters and conditioned on additional context, evolves alongside the student and its rollout context distribution. The benefit of modifying this moving target depends on how target–student probability mismatches translate into updates. We propose Self-Referenced On-Policy Self-Distillation (SR-OPSD), which constructs a normalized geometric target from the self-teacher and a frozen initial policy, then minimizes the forward R\'enyi divergence from this target to the student. The interpolation coefficient controls the self-teacher's contribution, while the R\'enyi order controls the power weighting of target-to-student probability ratios in the gradient. For fixed contexts and target components, we establish a conditional variational characterization and derive the exact token-logit gradient, revealing how anchoring and projection jointly shape the effective update target. Experiments across scientific reasoning, tool use, mathematical reasoning, and code generation demonstrate strong performance across multiple model families and scales. Ablations further show that reference anchoring can improve or degrade performance depending on the projection objective, supporting the joint design of target construction and projection geometry.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.