acceptodds
Under review as a conference paper at ICLR 2027

Learning Posterior Teachers for On-Policy Self-Distillation

Abstract

On-policy self-distillation addresses the sparse supervision of reinforcement learning with verifiable rewards (RLVR) by providing dense token-level supervision from a privileged teacher. However, privileged information alone does not ensure reliable supervision, and teacher accuracy alone does not specify a suitable distribution for distillation. We identify three desirable teacher properties for distillation: correctness, minimal shift, and preference preservation. The correctness-conditioned posterior unifies these properties by restricting the initial student's prior to trajectories with correct final answers. We introduce Flow Teacher Distillation (FlowTD), which trains a privileged teacher to approximate this posterior using GFlowNet trajectory balance. Its probability-matching objective uses outcome verification to supervise correctness and explicitly trains the teacher to preserve the student's relative probabilities among correct trajectories. Our analyses show that FlowTD learns teachers with stronger correctness discrimination than prompt-only teachers, alongside smaller distributional shifts and better preference preservation than SFT- and RL-trained teachers. Across four mathematical reasoning benchmarks, FlowTD improves average accuracy over the strongest baselines by 4.7 and 3.9 points on Qwen2.5-3B and Qwen3-8B-Base, respectively. It also achieves the highest pass@ among the compared teacher constructions on AIME 2025.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.