acceptodds
Under review as a conference paper at ICLR 2027

OPAD: On-Policy Action Distillation for Large Language Model Agents

Abstract

On-policy distillation trains language model agents with teacher guidance at student-visited states, but token-level matching can entangle decision learning with imitation of the teacher's text. The environment responds only to executed actions, making them a natural unit of knowledge transfer. We introduce On-Policy Action Distillation (OPAD), which aligns teacher and student preferences over executable actions. We show that the student's success gap to a reference that keeps the student's reasoning but executes the teacher's actions is bounded by missing or mismatched actions at student-visited states. OPAD builds teacher targets from sampled responses, scores candidate actions after the student's reasoning, concentrates supervision at local peaks of action-span entropy, and weights it by trajectory outcomes and teacher support to balance success consolidation and failure correction. The action interface supports cross-tokenizer and black-box teachers without vocabulary alignment or teacher logits. On ALFWorld, WebShop, and ScienceWorld, OPAD outperforms token-level and black-box distillation baselines in nearly all comparisons across student sizes, model families, and teacher access settings, with several students surpassing their teachers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.