The Gravity of Privilege: Tilting Learning Measures in On-Policy Self-Distillation
Abstract
On-Policy Self-Distillation (OPSD) is an emerging paradigm for LLM post-training, in which a model learns from a privileged view of itself through dense supervision on student-generated trajectories. However, existing OPSD methods incorporate privileged information through the conditional teacher distribution, while leaving the empirical measure over student-generated trajectories unchanged. This induces a structural asymmetry: privilege determines what the student learns at each visited trajectory, but not the measure under which different trajectories contribute to learning, even though privileged conditioning can vary substantially across sampled trajectories. To address this limitation, we propose Privileged Measure Self-Distillation (PMSD), which leverages privilege not only to construct distillation targets, but also to adaptively reweight the learning measure across training problems. Specifically, PMSD contrasts the same frozen teacher's support for each trajectory with and without privileged information, aggregating log-likelihood ratios into a policy-dependent signal for problem-level allocation. We formulate this adaptation as a relative-entropy trust-region optimization that assigns greater learning mass to problems receiving stronger support from privilege, while bounding the KL divergence from the original empirical measure. This mechanism intrinsically favors problems with larger estimated privileged likelihood contrasts while preventing extreme polarization in the learning measure. The resulting Gibbs-form solution is governed by a scalar dual variable, with the KL constraint directly bounding the induced shift in population gradient aggregation for the exact solution under bounded per-problem gradients. To obtain a more reliable estimate of the privilege-induced effect under the current policy, PMSD additionally uses multiple on-policy rollouts per problem. Extensive experiments on mathematical reasoning across Qwen3-1.7B/4B/8B show that PMSD consistently surpasses selected baselines by 0.84–2.50 percentage points in Avg@12 while maintaining superior token efficiency. Our code is available on https://anonymous.4open.science/r/pmsd-paper-code-B260.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.