CoEvoBench: Can LLMs Drive Recursive Co-Evolution of Agentic Systems?
Abstract
Autonomously improving agentic systems requires researchers to translate execution feedback into effective updates to model weights and the harness governing interaction with tools and environments. Existing evaluations primarily focus on model training or harness optimization, leaving research capability under joint access to both intervention types insufficiently assessed. We introduce CoEvoBench, a benchmark that evaluates LLM researchers on autonomous model–harness improvement across five agentic domains. Researchers start from a common initialization within each domain and independently choose training, harness revisions, and their combinations within a bounded resource budget, using development feedback to guide further research. The submitted model–harness pair then executes independently on held-out tasks to measure the system improvement delivered by the researcher. Across five frontier researcher models, we observe substantial but uneven gains across domains, with no researcher consistently achieving the strongest improvements on every benchmark. Research gains concentrate in key breakthroughs: among improving runs with at least three selection-set candidate evaluations, the largest increase in best-so-far performance accounts for a median of 62.5% of the total observed gain. These results highlight the challenge of extending isolated breakthroughs into sustained improvement. CoEvoBench provides a framework for evaluating this capability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.