acceptodds
Under review as a conference paper at ICLR 2027

Branching On-Policy Distillation

Abstract

On-policy distillation (OPD) provides teacher supervision on student-generated trajectories, but a single trajectory realizes only one continuation at each decision, limiting supervision coverage under a finite sampling budget. We propose BranchOPD (Branching On-Policy Distillation), which combines student uncertainty and teacher–student disagreement to select a few decision points, resamples continuations from shared prefixes under the student policy, and provides token-level teacher supervision on both main trajectories and branches. It supports token-level decisions in mathematics and within-turn decisions in agent tasks. We also introduce protocol-boundary repair (PBR), which uses targeted supervision under protocol constraints to mitigate repetition and wasted generation. BranchOPD outperforms Vanilla OPD and Multi-Rollout OPD () across all evaluated benchmarks, with gains over Vanilla OPD of 18.55 percentage points in mean mathematics accuracy over 16 sampled answers (avg@16) and 13.8 points in mean ALFWorld ID/OOD success. With PBR held fixed, prefix reuse reduces student-generated token costs on ALFWorld and WebShop by approximately 19% and 31%, respectively, relative to the seven-trajectory generation reference. Mechanistic analyses and selection ablations show that branching at uncertain decisions expands subsequent behavior coverage and outperforms random selection and turn-initial branching.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.