From Privilege to Parameters: Learning Dynamics of On-Policy Self-Distillation
Abstract
On-policy self-distillation (OPSD) has recently emerged as a promising post-training paradigm for improving language models. A teacher conditioned on privileged information, additional context available during training but unavailable at inference, supervises an unprivileged student's predictions at student-generated prefixes. However, how privileged information becomes internalized during training remains poorly understood. We revisit OPSD from the perspective of privilege transmission, combining a local learning-dynamics analysis with controlled random key–value recall experiments. We find that OPSD internalizes information through behavioral learning: privileged information is acquired only through the supervision it induces on the behavior the student is asked to produce. Internalization further exhibits a prefix-first progression aligned with the structure of autoregressive generation. These findings motivate **Recontextualizing Privileged Information**, a simple yet effective strategy that reuses privileged examples across task contexts while retaining task-local reference guidance, allowing information associated with one task to contribute to supervision on others. Experiments with Qwen3 models ranging from 1.7B to 8B parameters demonstrate improvements over vanilla OPSD, with gains of up to **17.61 percentage points** in best-checkpoint mean Avg@12 across eight mathematical reasoning benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.