acceptodds
Under review as a conference paper at ICLR 2027

Distilling Structure, Not Solutions: Structural Bottlenecks for On-Policy Self-Distillation

Abstract

On-policy self-distillation (OPSD) aligns supervision with student-generated trajectories but retains an information asymmetry: the teacher receives a verified solution that is unavailable to the student at inference. We investigate whether restricting this privileged context to reasoning structure improves distillation. **Structural-Bottleneck OPSD (SB-OPSD)** replaces full solutions with compact, result-free skeletons that preserve ordered reasoning operations while excluding intermediate results and final answers. A frozen, skeleton-conditioned teacher supplies next-token targets at student-generated prefixes, while the student uses a question-only prompt throughout training and inference. Across Qwen3-1.7B, 4B, and 8B on three competition-mathematics benchmarks, SB-OPSD achieves the highest mean **Average@12** among the compared methods, improving over the respective base models by **5.70, 3.31, and 3.22 percentage points**. Structural controls favor compact, ordered, question-matched skeletons. Relative to full-solution OPSD, SB-OPSD also reduces teacher-context and output lengths and exhibits a smaller gap to its conditioned teacher during training. These results support compact reasoning structure as an effective form of privileged supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.