Controlled In-Rollout Privileged Information Intervention for Reliable On-Policy Self-Distillation
Abstract
On-policy self-distillation (OPSD) trains a model on its own rollouts using token-level supervision from a privileged teacher instantiated from the same base model. In standard OPSD, the teacher receives the complete reference solution in its initial prompt. We study a specific failure mode of this design: full prompt-level privileged information (PI) can induce a concentrated teacher–student discrepancy near the beginning of a rollout, while its corrective influence weakens as an erroneous student prefix grows. We introduce Controlled In-Rollout Privileged-Information Intervention (CIPI), which periodically selects PI fragments by balancing their guidance benefits against the resulting teacher–student policy discrepancy, and inserts them into the teacher context along student-generated rollouts. Theoretical analysis provides conceptual motivation for balancing teacher guidance and teacher–student discrepancy in PI selection. Experiments with Qwen3-1.7B and Qwen3-4B on three competition-level mathematical reasoning benchmarks show that CIPI consistently outperforms OPSD, with macro Avg@12 gains of 3.3 and 5.1 percentage points (ppts) in non-thinking mode and 1.7 and 1.0 ppts in thinking mode, respectively. Ablation studies show that removing periodic in-rollout intervention while retaining prompt-level PI selection reduces average performance by 2.7 ppts, further confirming the effectiveness of in-rollout intervention beyond prompt-level PI selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.