CompEvolve: Competitive Evolution of LLM Agent Harness Regularized by Behavioral Diversity and Generalization Judging
Abstract
Evolutionary search is a well-established paradigm for automatic design, and has recently been applied to evolve the harness around a large language model (LLM): its prompt, control flow, and multi-agent structure. Because it modifies the harness rather than the model weights, this paradigm extends even to closed-source models, at lower cost and with greater interpretability than fine-tuning. Within it, competition, keeping only the fittest of competing harnesses, is a natural mechanism for directed improvement. We find, however, that naive competition among LLM agents does not reliably surpass independent, non-competitive evolution: the population rapidly converges to a few near-identical architectures, and selection by training accuracy alone rewards designs that overfit the training problems. We present CompEvolve, which turns competitive evolution into a reliable optimizer through three components: information isolation prevents a capable LLM architect from exploiting the training signal, compelling it to improve genuinely transferable ability; selection regularized by behavioral diversity retains agents that are initially weak but complementary to the rest of the population, which are more likely to be recombined into stronger agents in later generations; and an LLM generalization judge evaluates the transferability of an agent's design at the architecture level, complementing training accuracy. CompEvolve consistently improves out-of-distribution generalization over independent evolution across multiple reasoning benchmarks and LLMs of varying capability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.