acceptodds
Under review as a conference paper at ICLR 2027

WorldOPD: World-Knowledge Agentic On-Policy Distillation

Abstract

Agentic on-policy distillation (OPD) trains a student to match a teacher’s action distribution at states the student visits. Yet action distributions alone do not tell the student how the environment changes, how tools respond, or what consequences an action is likely to have. We call this information world knowledge (WK). Such knowledge may be especially hard for a smaller student to infer from multi-turn trajectories. We find that teacher-generated world knowledge reduces teacher–student divergence at most rollout states, but can increase it at some states. World knowledge has a state-dependent contextual effect, motivating selective auxiliary supervision. We introduce WorldOPD, which teaches students both how to act and what their actions may lead to. Alongside standard action OPD, WorldOPD distills the teacher's environmental knowledge through student-generated action–state pairs, supervising both candidate actions and their anticipated consequences. Since the benefit of world knowledge varies across interaction states, a lightweight predictor learns from sparse probes to select when to apply this additional supervision. Extensive experiments across function-calling and multi-turn search tasks show that WorldOPD consistently improves student performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.