Can Harness Optimization Improve Agents? Generalizable Gains Versus Evaluation Noise
Abstract
Harness optimization methods revise an agent’s prompts, tools, memory, and run- time control without updating model weights. We ask which of their improve- ments can be distinguished from evaluation noise, and where the reliable ones come from. We evaluate Meta-Harness, Agentic Harness Engineering, and Ret- rospective Harness Optimization on Terminal-Bench 2.1, DeepSWE, and Fron- tierCS, eight method–benchmark evaluations in all. In each, the selected harness and its seed harness receive repeated independent runs on every held-out task un- der the same model, budgets, and environment. The benchmark-level effect is an equally weighted mean over tasks, and its confidence interval reflects run-to-run variation. Two of the eight selected harnesses improve reliably on their seed har- nesses, five effects cannot be distinguished from evaluation noise, and one is a reliable regression. In an additional Meta-Harness search on Terminal-Bench 2.1, most of the selected harness’s held-out improvement comes from tasks on which the seed harness’s shell tends to crash. The selected harness restarts the shell after a crash; on the other tasks, there is no clear evidence of improvement. Repairing such an environment failure can raise held-out performance without showing that the agent solves problems better. We analyze change targets and observed failure causes to distinguish evidence of environment repair from evidence of improved problem solving.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.