acceptodds
Under review as a conference paper at ICLR 2027

Context-consistent Teacher Signals for On-policy Distillation

Abstract

On-policy distillation (OPD) provides dense teacher supervision on student- generated trajectories, but the teacher signal is typically evaluated under a single context. Such single-context supervision can be sensitive to contextual changes, making it difficult to distinguish stable learning signals from context-specific teacher preferences. Can we identify teacher supervision that remains consistent across task-relevant contexts and use it to improve on-policy distillation? We introduce Factorial On-Policy Distillation (FOPD), which evaluates identical student actions across teacher checkpoints and contexts to identify cross-context consistent supervision. FOPD centers the resulting teacher signals over the student’s action distribution, retains directions consistently supported across contexts with a conservative shared magnitude, and anchors the update to the standard OPD direction. On Qwen3-1.7B in non-thinking mode across 14 mathematics, code, and science benchmarks, FOPD achieves a Macro score of 43.35, corresponding to a 5.9% relative improvement over the strongest baseline, and improves performance on 13 of the 14 tasks. Ablation studies show that removing or replacing any defining component reduces Macro performance by 1.82–5.85 points. Controls matched for reward sparsity or magnitude remain 2.35–2.65 points below FOPD, while token-level evaluation shows that FOPD achieves the highest direction precision throughout the evaluated 10–70% coverage range. These results demonstrate that cross-context consistency provides an effective principle for identifying reliable token-level supervision in on-policy distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.