DOSD: Difficulty-privileged On-policy Self-Distillation
Abstract
On-policy self-distillation trains an unassisted student using a reference-conditioned teacher evaluated on the student's own trajectories. Reference access can supply useful supervision, but can also make that supervision depend on information unavailable to the student. We introduce Difficulty-privileged On-policy Self-Distillation (DOSD), which allocates teacher-side reference context using a per-problem rolling record of verified outcomes. An offline-initialized, fixed-length queue is updated with each current student response before the teacher scores that same response. Its success fraction determines the token fraction of a solution prefix exposed to the teacher, while the final answer remains separately available to the teacher and all privileged information remains hidden from the student. We formalize this training procedure, establish bounded changes in its exposure budget, and characterize the additional outcome-conditioned target variability introduced by immediate feedback. These properties motivate controlled tests of adaptive exposure; they do not establish a performance improvement or the elimination of privileged-information dependence. Across Qwen3-1.7B, 4B, and 8B, DOSD changes three-benchmark average accuracy relative to OPSD by , , and percentage points, respectively. In a Qwen3-8B-thinking controller comparison, DOSD exceeds a reported budget-calibrated fixed-prefix controller by 2.32 points and is 0.18 points below full-prefix OPSD. Because the available experiment record contains aggregate point estimates but not repeated seeds, complete run metadata, or realized context budgets, these results are descriptive rather than evidence of statistical significance or reduced end-to-end cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.