acceptodds
Under review as a conference paper at ICLR 2027

ForkPilot: Self-Evolving Policy for Retrospective Search in Long-Horizon Agents

Abstract

Interactive language-model agents increasingly solve complex tasks through long-horizon, multi-call reasoning, where errors in beliefs or actions can compound across tool interactions. Retrospective search can recover from such failures but is prone to misallocation. Delayed outcomes obscure the contribution of intermediate search decisions, leading to Attribution Complexity, while evolving execution evidence leads to Adaptation Complexity, where previously learned estimates become stale. To address these challenges, we first introduce Search Value Dynamics (\method), which characterizes the evolving trade-off between the gain and cost of retrospective search. Building on \method, we propose , a self-evolving two-stage policy-learning framework. In the first stage, learns a search-value policy offline from completed trajectories through automatically constructed outcome comparisons. In the second stage, it makes search decisions on the current observations and then self-evolves, incorporating newly completed trajectories into subsequent policy updates. We evaluate across 6 diverse benchmarks and 7 widely used LLM backbone families, including four open-source families, GPT-5.6 Sol and Opus 4.8, against 9 competitive baselines, including real-world harness deployment used by hundreds of thousands of paid users. achieves comparable state-of-the-art performance while reducing token usage by up to 59.2%, demonstrating its efficacy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.