Reference-Free On-policy Self-Distillation via Self-Reflection
Abstract
On-policy self-distillation (OPSD) provides dense, token-level supervision through a teacher policy conditioned on privileged information (PI), typically in the form of external reference solutions that require additional annotation or generation. Besides this acquisition cost, mismatches in reasoning paths and styles between these references and the student's own rollouts can impair student learning, motivating the use of PI that better aligns with the student's own reasoning. Reflection offers a natural source of such PI: it revisits a student's responses to provide feedback tailored to the reasoning they contain. We therefore investigate whether conditioning the teacher on self-reflections generated by the shared policy can improve on-policy self-distillation by providing more effective token-level supervision. We propose Reference-Free Self-Distillation via Self-Reflection (RF-SD), where a single policy serves as the student, reflector, and teacher. Given a question and a student rollout, the reflector is routed by the verifier outcome to summarize useful reasoning in correct trajectories or diagnose errors in incorrect ones. The resulting reflection replaces external references as the teacher's PI for token-level supervision on the student's own rollout. We further introduce RF-SD-e, which evolves the privileged information by optimizing the reflector using the correctness of reflection-conditioned teacher responses as rewards. Experiments on challenging mathematical reasoning benchmarks show that RF-SD outperforms strong baselines across different model scales, with further gains from reflector evolution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.