MEvo: Multi-Evolver Coordination for Self-Improving Agent Harnesses
Abstract
Self-improving agent harnesses in the DGM lineage all delegate each self-modification to a single LLM proposer—the single-editor assumption. We introduce MEvo (Multi-Evolver Coordination), to our knowledge the first harness-evolution configuration whose final write authority rests with an explicit protocol over k independent evolvers. We instantiate three protocols: MEvo-Debate (competing proposals under judge adjudication), MEvo-Pipeline (staged diagnose–modify–verify), and MEvo-MasterWorker (persistent cross-iteration authority). Across 5 seeds on the full 608-task EvoBench benchmark, Debate attains the highest Best macro, above the single-evolver baseline, while Pipeline holds the deployed score within 0.05 of its highest iterate—several-fold tighter than every other tested configuration. No protocol's final iterate surpasses the seed harness under the 3-iteration budget; the same incentive-disjoint discipline, lifted to the outer loop as a Preserve-and-Extend Gate, pushes Pipeline above seed at 5 iterations, and an Ensemble-of-Iterates recipe further improves inference-time quality. These results establish multi-evolver coordination with incentive-disjoint review as the pre-condition for converting transient exploration into net improvement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.