acceptodds
Under review as a conference paper at ICLR 2027

StyleOPD: Selecting the Generation Prompt for On-Policy Distillation

Abstract

On-policy distillation (OPD) supervises a student on its own rollouts. Recent work improves that supervision by changing what the teacher is conditioned on or which tokens receive loss, while the prompt under which the student generates its rollouts is left at the task default, even though those rollouts determine what supervision exists at all. We treat the generation prompt as a selection variable and choose it by measurement. Reusing the criterion behind teachability-based token selection (corrective disagreement the student can still act on), we turn a per-token gate into a per-prompt score, computed once on the untrained student with no labels, no training and no change to the rest of the OPD recipe. In a controlled study distilling Qwen3.5-27B into Qwen3.5-9B, the choice of prompt moves accuracy by more than the prompt-free OPD gain over the initial student, every added prompt improves on vanilla OPD averaged over training, and the selected prompt beats vanilla OPD on every seed. Disagreement is what decides: every disagreement-based signal we measured, including our score, selects the best prompt, while every rule built on agreement or accuracy selects a worse one, and the best training prompt is the one under which the untrained student is least accurate. The selected prompt outperforms a sequence-level distillation baseline that needs extra teacher data, and it composes with token selection: the prompt helps on every seed with or without token selection, and more with it over training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.