Opinionated Optimizers: Why Agents Struggle to Improve Agents
Abstract
Meta-optimization systems are increasingly used to improve agent harnesses for a given task through an iterative loop in which a meta-agent proposes edits, evaluates the resulting harnesses, and selects new candidates to build on. Prior work has shown that meta-optimized agents often fail to generalize beyond the data seen during optimization, but the reasons for this remain unclear. We argue that this overfitting occurs because the proposer meta-agent is opinionated: updates reflect prior preferences for particular strategies, unlike the analytically derived updates of gradient-based learning. We test this account through a systematic analysis of over 3 million recorded agent rollouts, organized around 216 optimization runs spanning 4 frontier methods and 3 benchmarks. To do so, we develop the Meta-Optimization Testbed (MOT), a standardized framework for studying meta-optimization in a controlled setting. Our analysis identifies 9 recurring failure modes spanning 3 stages of the meta-optimization loop: (a) proposals draw from a narrow set of familiar strategies shaped more by the proposer's priors than by the task, (b) noisy evaluations provide weak evidence for revising those priors, and (c) selection retains local improvements under method-specific rules, but further search rarely compounds them into additional gains. By tying failures to specific stages of the meta-optimization process, our analysis provides a clear roadmap for future research; we release MOT and the annotated traces to support it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.