acceptodds
Under review as a conference paper at ICLR 2027

Select Before You Act: Pre-Commitment Guidance for Agentic On-Policy Distillation

Abstract

On-policy distillation transfers the capabilities of language model agents through teacher supervision on student-generated interactions. In multi-turn tasks, however, student action errors can alter the environment and undermine the conditions for effective supervision at later turns. We introduce **PREACT**, a pre-commitment teacher-guided approach to multi-turn on-policy distillation. At each turn the student proposes several candidate responses, and a fixed teacher selects one for execution *before* its action reaches the environment Only the selected response is retained for distillation and appended to the interaction history with the resulting observation. Selecting a response before execution shapes not only the current supervision target but also the states in which subsequent learning occurs. **PREACT** thus extends teacher guidance from supervising individual responses to shaping the interactions that produce future training data, while keeping the student as the sole source of responses. Experiments on ALFWorld, ScienceWorld, and WebShop demonstrate that **PREACT** can improve student task performance across multiple model scales and under different initialization settings. These results highlight action commitment as an effective point of teacher guidance for training multi-turn agents. Our code is available at this [link](https://anonymous.4open.science/r/41099/).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.