acceptodds
Under review as a conference paper at ICLR 2027

From Privilege to Parameters: Learning Dynamics of On-Policy Self-Distillation

Abstract

On-policy self-distillation (OPSD) has recently emerged as a promising post-training paradigm for improving language models. A teacher conditioned on privileged information, additional context available during training but unavailable at inference, supervises an unprivileged student's predictions at student-generated prefixes. However, how privileged information becomes internalized during training remains poorly understood. We revisit OPSD from the perspective of privilege transmission, combining a local learning-dynamics analysis with controlled random key–value recall experiments. We find that OPSD internalizes information through behavioral learning: privileged information is acquired only through the supervision it induces on the behavior the student is asked to produce. Internalization further exhibits a prefix-first progression aligned with the structure of autoregressive generation. These findings motivate **Recontextualizing Privileged Information**, a simple yet effective strategy that reuses privileged examples across task contexts while retaining task-local reference guidance, allowing information associated with one task to contribute to supervision on others. Experiments with Qwen3 models ranging from 1.7B to 8B parameters demonstrate improvements over vanilla OPSD, with gains of up to **17.61 percentage points** in best-checkpoint mean Avg@12 across eight mathematical reasoning benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.