RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Abstract
Multi-turn agents trained with reinforcement learning (RL) from verifiable rewards receive one scalar reward per trajectory, which gives little guidance for intermediate decisions. On-policy distillation (OPD) adds dense token-level supervision from a teacher; in agentic training, the teacher is conditioned on privileged information, such as task skills available only during training, so that a skill-free student can internalize them. This recipe rests on two assumptions, and we find that neither holds in agentic tasks. First, privileged information alone does not make a teacher reliable: a skill-conditioned policy that is not itself optimized for the task often fails to outperform the student it supervises. Second, the benefit of teacher supervision is stage-dependent: matching the teacher accelerates early learning but later conflicts with reward optimization and holds the student near the teacher. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over GRPO by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and it surpasses its own skill-conditioned teacher in every setting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.