acceptodds
Under review as a conference paper at ICLR 2027

Learning to Localize and Repair Deviations in Long-Horizon Reasoning

Abstract

Large language model (LLM) agents increasingly tackle complex long-horizon tasks through multi-step reasoning, tool use, and environment interaction. However, local errors can propagate through subsequent steps, derailing reasoning trajectories and wasting computation. Existing approaches often provide scalar supervision with limited corrective guidance or rely on reflection and revision that incur additional inference overhead. We therefore propose **RLOR**, an observation-guided trajectory repair framework. RLOR identifies the earliest deviation as a pivot, injects observation-based natural-language feedback, and repairs only the subsequent continuation while preserving the valid prefix. We use multi-round pivot annotations to guide curriculum supervised fine-tuning (SFT) and pivot-aware reinforcement learning (RL) with suffix-localized credit assignment. This training enables agents to internalize and generalize repair behavior, allowing autonomous correction without external teacher feedback at deployment. Experiments on long-horizon reasoning benchmarks show that RLOR improves accuracy and token efficiency, with ablations confirming the contributions of pivot localization, observation content, and the training design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.