acceptodds
Under review as a conference paper at ICLR 2027

When Does Finer Teacher Routing Help? Per-Token Capability Routing in On-Policy Distillation

Abstract

On-policy distillation from several domain specialists raises a choice a fixed corpus never poses: which teacher supervises each token of the student's own rollout. The standard recipe routes a whole rollout to one teacher by its domain tag, yet when a rollout interleaves capabilities, that one label is wrong on part of the sequence. We propose (Capability-Routed On-Policy Distillation), which assigns a teacher to every token by vocabulary identity: offline, each token id goes to the specialist that most improves its next-token distribution over a shared initialization, and nothing else changes. With granularity isolated, the benefit is conditional. On data that mixes mathematics and code within each sample, per-token routing beats each single-teacher baseline on its weak axis, by points on code and on mathematics. The code gain rises monotonically with within-sample mixing, reverses to a -point loss on single-domain data, and is not detected in a B dense student or a B mixture-of-experts with B active parameters. Per-token routing can help only if the two single-teacher baselines split, a necessary condition that can be checked from runs a practitioner trains anyway, before the per-token router is built.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.