acceptodds
Under review as a conference paper at ICLR 2027

Flux-OPD: On-Policy Distillation with Evolving Contexts

Abstract

Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts derived from task feedback can provide richer supervision, yet fixed contexts cannot adapt as the student evolves, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in unstable distillation targets and conflicting distributions, requiring mechanisms to stabilize the target and downweight conflicts. In this paper, we characterize context distillation through a reverse-KL decomposition, showing that the student is distilled toward the geometric mean of context-conditioned teachers, while the objective contains a conflict term measuring disagreement among these teachers. Motivated by this characterization, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to incorporate contextual signals derived from task feedback. Flux-OPD treats differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher, and weights the correction strength using the conflict term. Experiments on open-ended tasks show that Flux-OPD achieves the best reported aggregate results among the evaluated OPD paradigms, highlighting the potential of combining teacher supervision with evolving contexts as in-training supervision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.