acceptodds
Under review as a conference paper at ICLR 2027

AgentReroute: Training LLMs to Correct Agent Trajectories with Reinforcement Learning

Abstract

Once a large language model agent makes an erroneous decision early in a task, the resulting error can accumulate along its autoregressive trajectory and ultimately lead to task failure. Dominant self-correction methods rely on reflection and feedback to improve outputs, but the model's self-evaluation is often influenced by its original erroneous judgments, making it difficult to correct accumulated errors. To this end, external assistance methods introduce dedicated modules for diagnosis and feedback to correct accumulated errors. However, these methods often rely on heuristic rules or prompt-based judgments, making it difficult to achieve both accurate rollback and effective correction. To address this issue, we propose AgentReroute, an externally assisted reinforcement learning framework that uses task outcomes as reward signals to jointly optimize rollback control and corrective feedback. Specifically, AgentReroute freezes the action generator, Generator, and trains only an independent recovery controller, Rerouter. Given the task instruction, the active trajectory, and the recovery record, Rerouter decides whether to continue execution or select a rollback point and generate corrective feedback, thereby guiding Generator to regenerate subsequent actions. Extensive experiments on four question-answering tasks, WebShop, and the Game24 interactive environment show that AgentReroute significantly improves task success rates while keeping Generator frozen, and that it can be combined with Reflexion for further gains. Further analysis shows that the learned recovery policy reduces repetitive behavior after rollback and maintains stable gains when reused across question-answering datasets and across Generators, demonstrating strong generalization and cross-model transferability. Our code and dataset are anonymously available at https://anonymous.4open.science/r/AgentReroute-DCCA/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.