Flux-OPD: On-Policy Distillation with Evolving Contexts
Abstract
Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts derived from task feedback can provide richer supervision, yet fixed contexts cannot adapt as the student evolves, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in unstable distillation targets and conflicting distributions, requiring mechanisms to stabilize the target and downweight conflicts. In this paper, we characterize context distillation through a reverse-KL decomposition, showing that the student is distilled toward the geometric mean of context-conditioned teachers, while the objective contains a conflict term measuring disagreement among these teachers. Motivated by this characterization, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to incorporate contextual signals derived from task feedback. Flux-OPD treats differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher, and weights the correction strength using the conflict term. Experiments on open-ended tasks show that Flux-OPD achieves the best reported aggregate results among the evaluated OPD paradigms, highlighting the potential of combining teacher supervision with evolving contexts as in-training supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.