acceptodds
Under review as a conference paper at ICLR 2027

SelfSearch: Reward-Free Search for Self-Improving Agents

Abstract

Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce SelfSearch, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch discovers agents that improve success rates by up to 11.2 percentage points on Terminal-Bench 2.1, 6.7 on SWE-bench Multilingual, and 5.0 on SWE-bench Verified. On SWE-bench Multilingual, an agent improves success by 5.0 percentage points while reducing execution cost by 38.5% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only $4.03 in search cost, it produces a harness achieving 82.0% on Terminal-Bench 2.1 using DeepSeek V4 Flash, comparable to the reported performance of Codex, the top-performing harness in a published nine-harness comparison. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.