Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?
Abstract
Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop. We ask whether this task-specific procedure remains necessary when a frontier model acts as the optimizer. We introduce Open-Ended Optimization (OEO), which fixes the objective, permitted interactions, resource budget, data boundary, and evaluation while allowing the optimizer to compose the improvement process online. We compare OEO with two complementary prescribed approaches: SkillOpt, a staged pipeline with bounded edits, and GEPA, a reflective evolutionary search. Across a complete grid of 16 head-to-head comparisons over 8 benchmark–target-model settings, GPT-5.5-driven OEO produces 12 higher final-score point estimates, 1 tie, and 3 lower estimates, with a median absolute difference of 3.1 percentage points. It improves the shared initial skill in all 8 settings while using a median 34.3% of SkillOpt’s configured target-interaction token allocation. Across 18 repeated optimization runs on 2 anchor settings, OEO retains the highest mean on LiveMath, whereas the 3 methods’ SearchQA means lie within 0.84 percentage points. A capability probe further exposes the boundary of delegation: a weak optimizer cannot execute the open-ended loop under the fixed interface, prescription is task-dependent at medium capability, and frontier OEO is competitive on both tested tasks. A one-shot control shows that the gains are not explained by a single prior-driven rewrite. Finally, the fully instrumented OEO–SkillOpt pair shows that prescription changes committed optimization paths more consistently than selected item-level behavior. Together, these results suggest that task-specific optimization procedures need not always be prescribed in advance. With a sufficiently capable optimizer, the procedure itself can instead be composed online within externally fixed objectives, permitted interactions, resource budgets, data boundaries, and evaluation, while explicit prescription remains a task- and capability-dependent scaffold rather than a universal requirement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.