Context-consistent Teacher Signals for On-policy Distillation
Abstract
On-policy distillation (OPD) provides dense teacher supervision on student- generated trajectories, but the teacher signal is typically evaluated under a single context. Such single-context supervision can be sensitive to contextual changes, making it difficult to distinguish stable learning signals from context-specific teacher preferences. Can we identify teacher supervision that remains consistent across task-relevant contexts and use it to improve on-policy distillation? We introduce Factorial On-Policy Distillation (FOPD), which evaluates identical student actions across teacher checkpoints and contexts to identify cross-context consistent supervision. FOPD centers the resulting teacher signals over the student’s action distribution, retains directions consistently supported across contexts with a conservative shared magnitude, and anchors the update to the standard OPD direction. On Qwen3-1.7B in non-thinking mode across 14 mathematics, code, and science benchmarks, FOPD achieves a Macro score of 43.35, corresponding to a 5.9% relative improvement over the strongest baseline, and improves performance on 13 of the 14 tasks. Ablation studies show that removing or replacing any defining component reduces Macro performance by 1.82–5.85 points. Controls matched for reward sparsity or magnitude remain 2.35–2.65 points below FOPD, while token-level evaluation shows that FOPD achieves the highest direction precision throughout the evaluated 10–70% coverage range. These results demonstrate that cross-context consistency provides an effective principle for identifying reliable token-level supervision in on-policy distillation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.