SAGE-OPD: Selective Agent-Guided Intervention for Multi-Turn On-Policy Distillation
Abstract
On-policy distillation (OPD) improves student models by training them on trajectories induced by their own policy, making it a promising approach for mitigating exposure bias in agent training. However, most OPD studies focus on single-turn settings, while realistic LLM agents interact with environments over multiple turns. In this regime, early errors can alter future observations and compound across the trajectory, and standard dense token-level OPD becomes brittle, as it may over-penalize semantically valid alternatives, reinforce local degeneracies such as repeated actions, and propagate unreliable teacher supervision on off-distribution histories. We propose SAGE-OPD, a verifier-free selective intervention framework specifically designed for multi-turn OPD. Instead of applying teacher supervision uniformly, SAGE-OPD uses pre-execution validity checks and teacher judgment to decide whether each student response requires supervision. SAGE-OPD further weights the selected turns by teacher confidence, assigning greater relative weight to turns with more decisive teacher predictions. Finally, normalizing weights within each batch preserves the relative allocation of supervision while maintaining unit average effective token weight. With a Qwen3-1.7B student and Qwen3-8B teacher, SAGE-OPD raises ALFWorld unseen success from 61.94% to 70.15% and ScienceWorld success from 4.70% to 6.71% over standard OPD, while remaining competitive on SearchQA. Component ablations isolate intervention and confidence weighting; three-seed replications, expert-action comparisons, and shared-state continuation evaluations provide complementary evidence of robustness and agent behavior. These findings support feedback-conditioned turn-level supervision allocation for multi-turn OPD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.