Rewind-OPD: Discrepancy-Guided Trajectory Rollback for On-Policy Distillation in Agentic Tasks
Abstract
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories. In agentic tasks, an early divergent action can change subsequent observations, allowing teacher–student discrepancies to accumulate across turns. Existing intervention methods typically act after this drift has become apparent, when the current state may already reflect an earlier mistake. We propose Rewind-OPD, which uses the history of teacher–student discrepancies to select an earlier intervention point, then rolls back to that state and reconstructs the trajectory under teacher guidance before returning control to the student. Across three teacher–student pairs and three agentic benchmarks, Rewind-OPD achieves the best performance among six baselines. With a general 30B teacher, it improves success rate over standard OPD by 4.14 and 4.66 percentage points for 1.7B and 4B students, while reducing average output tokens by 16.0% and 8.0%, respectively. With specialized 8B teachers, the 4B student exceeds the strongest baseline and teachers in average success rate by 1.72 and 1.49 percentage points, respectively, supporting its effectiveness and generalizability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.