acceptodds
Under review as a conference paper at ICLR 2027

DarwinX: Evolving Agent Harnesses through Natural Selection

Abstract

An LLM agent’s capabilities depend on the harness that supplies its instructions, tools, memory, and control flow. Improving that harness is a selection problem: an edit that fixes one task can regress another, a discarded variant may contain a complementary capability, and a noisy local gain may not repeat. We introduce DarwinX, a population-based method for evolving harnesses with the model fixed. Its preserve-and-extend contract screens task-level regressions, while independent confirmation determines which variants may steer later search. The archive retains other specialists for recombination, and execution traces and shared memory guide subsequent proposals. Search uses verifier-backed synthetic tasks that reuse benchmark environments; original benchmark tasks remain outside the search loop. On complete official evaluations, the selected harnesses improve five-attempt average success by 7.19 percentage points on Terminal-Bench 2.1 and 9.38 points on DeepSWE. On WebArena-Infinity, the frozen winner achieves 91.19% single-attempt success across 1,260 tasks, versus 75.48% for the composite report baseline, with no static-invalid evolved trajectory flagged by the action audit. These results show that the complete selection procedure can improve a frozen-model agent beyond the synthetic tasks used for search.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.