WorldOPD: World-Knowledge Agentic On-Policy Distillation
Abstract
Agentic on-policy distillation (OPD) trains a student to match a teacher’s action distribution at states the student visits. Yet action distributions alone do not tell the student how the environment changes, how tools respond, or what consequences an action is likely to have. We call this information world knowledge (WK). Such knowledge may be especially hard for a smaller student to infer from multi-turn trajectories. We find that teacher-generated world knowledge reduces teacher–student divergence at most rollout states, but can increase it at some states. World knowledge has a state-dependent contextual effect, motivating selective auxiliary supervision. We introduce WorldOPD, which teaches students both how to act and what their actions may lead to. Alongside standard action OPD, WorldOPD distills the teacher's environmental knowledge through student-generated action–state pairs, supervising both candidate actions and their anticipated consequences. Since the benefit of world knowledge varies across interaction states, a lightweight predictor learns from sparse probes to select when to apply this additional supervision. Extensive experiments across function-calling and multi-turn search tasks show that WorldOPD consistently improves student performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.