acceptodds
Under review as a conference paper at ICLR 2027

Lookahead-Guided Recursive Self-Improvement for Language Model Agents

Abstract

Recursive self-improvement can reduce the manual effort needed to build capable language model agents, but edits that help the current agent can become harmful after later modifications. We aim to select currently beneficial edits that retain their value as the agent continues to change. We introduce Lookahead-Guided Recursive Self-Improvement (LRSI), which measures each candidate's contribution in sampled future agent versions by enabling and disabling it while holding subsequent modifications fixed. LRSI selects the largest immediate improvement among candidates whose average contribution in the least favorable sampled continuations is nonnegative, keeping the language model fixed and committing only the current edit. Across ALFWorld, Terminal-Bench 2.0, SWE-bench Verified, and Aider Polyglot with three model backbones and matched optimization budgets, LRSI improves task success, pass, or resolved rates by 2.39–4.39 percentage points over the strongest evaluated baseline in each setting. By distinguishing an edit's continued usefulness from the overall quality of future agents, this work provides a testable criterion for selecting improvements that better withstand subsequent self-modification.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.