acceptodds
Under review as a conference paper at ICLR 2027

Divergence-Calibrated Turn Importance for Multi-Turn Agentic On-Policy Distillation

Abstract

On-policy distillation (OPD) has become a prevailing recipe for transferring multi-turn agentic capabilities from large teachers to small deployable students. However, effective multi-turn distillation remains challenging: OPD supervises at the token level, whereas an agentic decision is made by a complete turn, so token-level supervision can flatten turn-level decisions. Existing methods have recognized the multi-turn agentic property, but they either introduce environment rewards that pure OPD lacks or aim to skip unnecessary supervision, rebalance the loss budget, and keep supervision reliable. This leaves an open question: how should supervision be allocated across turns by their importance when no environment reward is available? We propose Divergence-Calibrated Turn Importance for Multi-Turn Agentic On-Policy Distillation (DT-OPD) to answer this question. We demonstrate that, under natural-gradient updates, the teacher-student divergence drift induced by a sampled response is determined by the response's stop-gradient negated divergence score and by its relative divergence. Our analysis lifts the divergence from the token level to the turn level, aligning it with the unit at which the agent acts. DT-OPD therefore turns the drift into a practical turn-level teacher-student divergence proxy and redistributes supervision across turns accordingly. Extensive experiments on ALFWorld, Search-QA, and WebShop, distilling Qwen3 and Qwen2.5 students from 0.6B to 4B using GRPO teachers up to Qwen3-30B-A3B, show that DT-OPD attains the best average in each benchmark at every student scale, and ablation studies and training dynamics further confirm that both of our core contributions are essential to this performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.