acceptodds
Under review as a conference paper at ICLR 2027

Controlling Interaction Drift: Optimizing LLM Agents for Multi-Turn User Interaction

Abstract

Large language models are increasingly post-trained as multi-turn agents for tool-mediated customer service, such as e-commerce assistance. In such settings, rollouts are generated online through a branching interaction process in which agent actions, user responses, and environment updates jointly determine what happens next. An erroneous response or tool action can therefore redirect subsequent interactions onto a different branch. We call this error-driven redirection *interaction drift*. This creates a collection-time decision: whether an error-conditioned branch should continue generating interactions for policy optimization. On-policy distillation (OPD) provides fine-grained teacher supervision in student-visited contexts, whereas terminal-reward reinforcement learning provides task-level feedback only after rollout completion. Neither signal determines whether another interaction should be collected from such a branch or how task credit should be localized within the trajectory. We introduce BLADE (Boundary-Localized Advantage for Drifted Episodes), a post-training framework with a frozen, environment-grounded Boundary Judge. After each completed nonterminal interaction loop and before the next simulated-user message, the Judge either continues collection or stops the branch at the loop in which the error was introduced. Teacher supervision is retained over the full collected rollout in either case, while the stopped-rollout task signal is confined to that loop; naturally completed rollouts retain trajectory-level task credit. Experiments across two stateful interaction benchmarks and two teacher–student configurations, including held-out-domain evaluation, show consistent improvements in strict repeated-run reliability over strong post-training baselines. These results suggest that controlling interaction drift during collection can improve the consistency of successful multi-turn behavior across model scales and evaluation settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.