FarOPD: Learning from Future Interactions in On-Policy Distillation for Multi-Turn Agents
Abstract
On-policy distillation (OPD) provides dense teacher supervision for training large language model (LLM) agents on student-generated trajectories. However, each action shapes future observations and contexts, making next-token optimization alone myopic in multi-turn settings: actions with similar local supervision can lead to substantially different future interactions. We show that the trajectory-level reverse-KL objective intrinsically connects an action's supervision to its influence on future interactions. Teacher-student mismatch thus provides not only a local matching signal, but also hindsight credit for earlier actions. Based on this insight, we propose FarOPD (On-Policy Distillation with Future Action Rewards), which combines local-token supervision with a future-action reward derived from subsequent interactions. This makes action updates sensitive to future teacher-student mismatch, extending local imitation toward trajectory-level matching. Extensive experiments on three multi-turn agent benchmarks demonstrate consistent improvements over OPD baselines, with success rate gains of up to 8.4% relative to Vanilla OPD. These gains extend to multi-teacher distillation and are especially pronounced on tasks requiring longer interactions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.