Do We Need Online Environment Interaction in Agentic OPD? An Empirical Analysis
Abstract
On-policy distillation (OPD) has become a common post-training approach where the student rolls out complete trajectories on its own, and the teacher provides token-level supervision. In multi-turn agentic tasks, on-policy rollouts of the student rely on iterative environment interaction to receive tool call feedback, incurring heavy time cost for long-horizon tasks. Yet, it remains poorly understood whether continuous environment interaction is necessary for agentic OPD. In this work, we analyze the limitations of agentic OPD with full rollouts, compare it with its environment-free alternative, and characterize the trade-offs between the two. We observe that in long-horizon task, teacher-student agreement increases with trajectory depth, yielding little supervision signal despite a substantial gap in task performance. In contrast, conditioning the student on replayed teacher prefixes retain stronger signal at later turns. Empirically, OPD with replayed prefixes matches or exceeds full-rollout performance on SWE-bench Verified and BFCL v4 with lower wall-clock cost, without online environment interaction. Our analysis reveals that OPD with replayed prefix achieve larger gains at pass@1 under a fixed token budget, whereas full rollouts brings pass@ closer to the teacher. Replayed prefixes reduce training instability associated with overlong responses under limited SFT cold start and yield larger gains across different cold-start levels, although more complete SFT initialization still leads to higher final performance. Together, our results establish that OPD with replayed prefix is ideal when training efficiency is the priority, while full rollouts are better suited for pushing the model's performance upper bound.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.