acceptodds
Under review as a conference paper at ICLR 2027

Stateful Search with Stateless Agents: Scaling Long-Horizon Automated Research

Abstract

Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay ever-growing histories, duplicate one another's work, or stop experimenting while token consumption continues. Yet most evaluations use short budgets or benchmarks that saturate early, leaving these failure modes untested. We introduce Agentic Evolutionary Search (AES), built on the principle of stateful search with stateless agents: the harness owns candidate solutions and measured outcomes, and reconstructs a fresh, role-specific context for every Advisor epoch and Worker trial. The Advisor turns accumulated evidence into concrete assignments for parallel Workers. We evaluate AES against three recent frameworks on software engineering, kernel optimization, and algorithm design at budgets of up to one billion cumulative tokens. AES achieves the best final result on every evaluated task and reaches the strongest kernel baseline's final performance with over 84% fewer tokens. Ablations from shared checkpoints show that focused contexts and explicit assignments each contribute, with effects that compound over full runs, while the Advisor consumes less than 0.6% of total tokens. These results argue for keeping durable research state out of agent conversations, and show that short evaluation horizons can misjudge both research systems and their components.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.