acceptodds
Under review as a conference paper at ICLR 2027

Move Selection Upstream: Choosing What the Teacher Trains On in On-Policy Distillation

Abstract

On-policy distillation (OPD) trains a student on its own generated responses under teacher supervision, and existing data-selection methods choose the student's prompts for a given teacher. Yet even a well-trained teacher performs unevenly across prompts, and selecting student prompts leaves it unchanged. When the teacher can also be trained, can selecting its training data produce a better final student under the same training budget? We select teacher-training data that targets the initial student's weaknesses, prioritizing data that is challenging for the initial student but demonstrably solvable. Our theoretical analysis gives sufficient conditions for teacher-data selection to outperform any selection of the student's prompts at the same training budget: the teacher must improve sufficiently and the student must train long enough. Empirically, teacher-data selection improves final-student performance in instruction following and mathematics, with absolute gains of 6.9% and 4.1% over uniformly sampled teacher-training data. It also outperforms the evaluated student-prompt selection baselines under matched training budgets. In replicated multi-teacher experiments, selecting the instruction-following teacher's data yields absolute gains of 2.2% to 5.2% over uniform sampling and student-prompt selection baselines on two public instruction-following benchmarks. Improving the teacher through data selection can thus be more effective than selecting the student's prompts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.