EvoG: A Modular Framework for Evolutionary Optimization of Agentic Workflows
Abstract
Designing effective agentic workflows for large language model (LLM) applications requires repeated, costly revision and evaluation. Evolutionary optimization can automate this process and existing methods have instantiated specific combinations of candidate selection, revision, example sampling, and evaluation strategies. Making these choices explicit and composable would help users construct and compare configurations suited to their task objectives and budget. Therefore, we introduce EvoG (Evolve on Graph), a modular framework for composing evolution strategies within a shared optimization loop. EvoG separates evolution into four components: a Selector that chooses parent candidates; a Generator that produces revisions; a Sampler that selects evaluation examples; and an Evaluator that assesses candidates. A shared history links candidate lineage, per-example outcomes, and execution traces to inform subsequent iterations. We evaluate EvoG on HotpotQA, MuSiQue-Ans, BigCodeBench-Hard, ScienceWorld, and WebShop. It achieves the highest mean test scores among the evaluated methods on four of five benchmarks, improving over the strongest baseline by 1.10 points on HotpotQA, 3.76 points on MuSiQue-Ans, 7.00 points on ScienceWorld, and 4.89 points on WebShop. On BigCodeBench-Hard, the best EvoG configuration achieves 35.33 points versus AgentSquare's 37.00 points, at approximately one-eighth its search cost and half its inference cost. For reproducibility, we release our code at https://anonymous.4open.science/r/EvoG-D6C2/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.