acceptodds
Under review as a conference paper at ICLR 2027

Controlled In-Rollout Privileged Information Intervention for Reliable On-Policy Self-Distillation

Abstract

On-policy self-distillation (OPSD) trains a model on its own rollouts using token-level supervision from a privileged teacher instantiated from the same base model. In standard OPSD, the teacher receives the complete reference solution in its initial prompt. We study a specific failure mode of this design: full prompt-level privileged information (PI) can induce a concentrated teacher–student discrepancy near the beginning of a rollout, while its corrective influence weakens as an erroneous student prefix grows. We introduce Controlled In-Rollout Privileged-Information Intervention (CIPI), which periodically selects PI fragments by balancing their guidance benefits against the resulting teacher–student policy discrepancy, and inserts them into the teacher context along student-generated rollouts. Theoretical analysis provides conceptual motivation for balancing teacher guidance and teacher–student discrepancy in PI selection. Experiments with Qwen3-1.7B and Qwen3-4B on three competition-level mathematical reasoning benchmarks show that CIPI consistently outperforms OPSD, with macro Avg@12 gains of 3.3 and 5.1 percentage points (ppts) in non-thinking mode and 1.7 and 1.0 ppts in thinking mode, respectively. Ablation studies show that removing periodic in-rollout intervention while retaining prompt-level PI selection reduces average performance by 2.7 ppts, further confirming the effectiveness of in-rollout intervention beyond prompt-level PI selection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.