Rethinking Agent Evolution: Do Evolved GUI Harnesses Actually Outperform Direct Inference?
Abstract
Program evolution is increasingly used to improve language-model agents, with progress typically measured against the initial program. We study whether seed-relative gains translate into an advantage over well-specified direct interfaces to the same frozen model. We introduce EvoHarness, a framework for evolving GUI grounding harnesses and auditing the resulting gains with frozen interface controls, matched execution backends, and repeated measurement. EvoHarness searches over executable preprocessing, prompting, parsing, and routing code, retaining execution traces for post hoc analysis. In a pilot search, evolution finds a harness that reaches 89.15% accuracy on all 1,272 ScreenSpot examples. Under matched backends, however, the evolved champion does not consistently outperform direct inference controls on the same frozen model: on Qwen, it improves from 20.33% to 54.33% over the seed but remains below a single-call direct-point baseline at 70.67% (paired Holm p=1.26×10^-5); on Kimi, it is not significantly better than either direct control. Repeated evaluation of identical code also yields apparent gains: a gain-only rule accepts 5/20 disjoint same-code pairs, while a paired gate accepts 0/20. These results suggest that seed-relative gains alone are insufficient evidence for agent improvement. We recommend frozen direct-interface controls, matched execution resources, and uncertainty reported at both the program and search levels.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.