acceptodds
Under review as a conference paper at ICLR 2027

CoEvoBench: Can LLMs Drive Recursive Co-Evolution of Agentic Systems?

Abstract

Autonomously improving agentic systems requires researchers to translate execution feedback into effective updates to model weights and the harness governing interaction with tools and environments. Existing evaluations primarily focus on model training or harness optimization, leaving research capability under joint access to both intervention types insufficiently assessed. We introduce CoEvoBench, a benchmark that evaluates LLM researchers on autonomous model–harness improvement across five agentic domains. Researchers start from a common initialization within each domain and independently choose training, harness revisions, and their combinations within a bounded resource budget, using development feedback to guide further research. The submitted model–harness pair then executes independently on held-out tasks to measure the system improvement delivered by the researcher. Across five frontier researcher models, we observe substantial but uneven gains across domains, with no researcher consistently achieving the strongest improvements on every benchmark. Research gains concentrate in key breakthroughs: among improving runs with at least three selection-set candidate evaluations, the largest increase in best-so-far performance accounts for a median of 62.5% of the total observed gain. These results highlight the challenge of extending isolated breakthroughs into sustained improvement. CoEvoBench provides a framework for evaluating this capability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.