acceptodds
Under review as a conference paper at ICLR 2027

SAGE-OPD: Selective Agent-Guided Intervention for Multi-Turn On-Policy Distillation

Abstract

On-policy distillation (OPD) improves student models by training them on trajectories induced by their own policy, making it a promising approach for mitigating exposure bias in agent training. However, most OPD studies focus on single-turn settings, while realistic LLM agents interact with environments over multiple turns. In this regime, early errors can alter future observations and compound across the trajectory, and standard dense token-level OPD becomes brittle, as it may over-penalize semantically valid alternatives, reinforce local degeneracies such as repeated actions, and propagate unreliable teacher supervision on off-distribution histories. We propose SAGE-OPD, a verifier-free selective intervention framework specifically designed for multi-turn OPD. Instead of applying teacher supervision uniformly, SAGE-OPD uses pre-execution validity checks and teacher judgment to decide whether each student response requires supervision. SAGE-OPD further weights the selected turns by teacher confidence, assigning greater relative weight to turns with more decisive teacher predictions. Finally, normalizing weights within each batch preserves the relative allocation of supervision while maintaining unit average effective token weight. With a Qwen3-1.7B student and Qwen3-8B teacher, SAGE-OPD raises ALFWorld unseen success from 61.94% to 70.15% and ScienceWorld success from 4.70% to 6.71% over standard OPD, while remaining competitive on SearchQA. Component ablations isolate intervention and confidence weighting; three-seed replications, expert-action comparisons, and shared-state continuation evaluations provide complementary evidence of robustness and agent behavior. These findings support feedback-conditioned turn-level supervision allocation for multi-turn OPD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.