acceptodds
Under review as a conference paper at ICLR 2027

Rewind-OPD: Discrepancy-Guided Trajectory Rollback for On-Policy Distillation in Agentic Tasks

Abstract

On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories. In agentic tasks, an early divergent action can change subsequent observations, allowing teacher–student discrepancies to accumulate across turns. Existing intervention methods typically act after this drift has become apparent, when the current state may already reflect an earlier mistake. We propose Rewind-OPD, which uses the history of teacher–student discrepancies to select an earlier intervention point, then rolls back to that state and reconstructs the trajectory under teacher guidance before returning control to the student. Across three teacher–student pairs and three agentic benchmarks, Rewind-OPD achieves the best performance among six baselines. With a general 30B teacher, it improves success rate over standard OPD by 4.14 and 4.66 percentage points for 1.7B and 4B students, while reducing average output tokens by 16.0% and 8.0%, respectively. With specialized 8B teachers, the 4B student exceeds the strongest baseline and teachers in average success rate by 1.72 and 1.49 percentage points, respectively, supporting its effectiveness and generalizability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.