Agentic On-Policy Distillation: Distilling Decisions, Not Tokens
Abstract
Distilling capable language-model agents into smaller models requires aligning the student's behavior with the teacher's. This raises a fundamental question: at which states and at what granularity should this alignment be performed? Offline imitation provides dense but token-specific supervision and can suffer from state-distribution mismatch. Token-level on-policy distillation (OPD) addresses this mismatch but typically requires teacher token probabilities and compatible tokenization. We argue that semantic actions provide a natural unit of supervision for agent distillation. We introduce Agentic On-Policy Distillation (AOPD), motivated by maximizing expected teacher-action likelihood at student-encountered states. At each such state, AOPD samples actions from both the student and a black-box teacher, rewarding student actions according to their semantic agreement with the sampled teacher actions. Through policy optimization, AOPD aligns the student's action-level behavior with the teacher's decisions without requiring reproduction of teacher-written reasoning or access to teacher token probabilities. Training Qwen3-4B and Qwen3-8B with two teachers on diverse, curated tool-use tasks yields average gains of 3.4-6.6 percentage points across four out-of-distribution benchmarks. AOPD scales favorably with student size and teacher capability, without requiring the teacher and student to share a model family or tokenizer. On -bench, AOPD achieves competitive performance against offline distillation and token-level OPD. We further explore Self-AOPD with a same-model teacher and privileged context. Together, these results establish AOPD as a promising approach to agent distillation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.