Attune: State-Adaptive Supervision Allocation for On-Policy Distillation
Abstract
On-policy distillation trains compact language models with teacher feedback on their own responses. Deciding which feedback to emphasize remains challenging: teacher–student disagreement signals a prediction gap, while the benefit of correcting it depends on the student's capacity and current training state. We introduce a controlled paired-intervention protocol using supervised fine-tuning on compositional arithmetic to measure the returns of fixed supervision under matched token exposure. The relative gains vary across student capacities and training states, supporting a learner-conditioned view of supervision utility. We operationalize this view through Attune, a coupled token–group allocation framework for on-policy distillation. Applying the same token calibration to candidate and reference losses couples the learning signal to the gradient comparisons used for allocation. Across eight benchmarks and two model families, Attune achieves higher aggregate scores than the compared distillation methods, and these gains persist across three training seeds. Independent-probe interventions further show that both components reduce distillation loss relative to uniform and shuffled weighting. Attune requires no difficulty labels and adds no inference-time computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.