RelEvolve: Learning What to Evolve, Evolving What to Learn
Abstract
Automated harness optimization improves LLM agents by iteratively revising the programs that govern prompting, memory, tool use, and control flow. Existing harness optimizers use prior harnesses and evaluation feedback to decide how to modify the next candidate, but the unresolved improvement directions that should guide subsequent evolution typically remain implicit. We introduce RelEvolve, which represents these directions as inquiries and couples inquiry updating with harness evolution: inquiries guide which candidate to continue from and what change to make, while the resulting evaluation provides evidence to refine the inquiry state for later evolution rounds. RelEvolve realizes this coupling through a Harness Lineage that records executable inheritance, an Inquiry Graph that maintains inquiry state across branches, and a Relational Trace that links the inquiries guiding each evolution to the resulting harness change and evaluation evidence. Across five benchmarks and two LLM backbones, RelEvolve yields a 4.4% absolute improvement in held-out test performance over the strongest evaluated baseline in each setting and a 12.5% absolute improvement over the unoptimized harness, averaged across settings. On HotpotQA, it reaches 89.84 and 91.24 F1 with GPT-5.6 Luna and GLM-5.3-Flash, respectively. Three independent search seeds on HotpotQA and InterCode-SQL maintain this improvement in these two representative settings, and on HotpotQA, ablations further indicate that weakening the coupling between inquiry updating and harness evolution lowers performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.