acceptodds
Under review as a conference paper at ICLR 2027

WHAT IS ONE SEARCH OPERATION WORTH? EVALUATING ITS DOWNSTREAM VALUE IN AUTOMATIC AGENT DESIGN

Abstract

Automatic agent design searches for better prompts, agent connections, and workflows. Yet a strong final system does not reveal which individual search decisions improved it. We introduce a protocol for evaluating a search operation through the optimization that follows it. From the same recorded state, we execute the original operation or a specified alternative, then resume optimization under the same objective and budget. We apply this protocol to GPTSwarm, AFlow, MaAS, and EvoMAS across six benchmarks. A dedicated GPTSwarm–MMLU study includes 480 executions at 20 states, with same-operation controls and separate runs for selection and validation. Operations selected using four repetitions per alternative show an apparent mean advantage of 2.17 percentage points over the unselected alternatives. Four additional repetitions on the same evaluation items estimate this advantage at 0.65 points (95% CI ), yielding a paired selection–validation gap of 1.52 points . GPTSwarm and EvoMAS/MMLU also show measurable separation in subsequent search trajectories, while estimated associations with downstream value remain imprecise. These measurements distinguish how an operation changes search from how well a selected operation performs on new executions, enabling evaluation of individual decisions within automatic agent design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.