TriEvo: Three-Agent Evolution of Policies and Optimizers
Abstract
An adaptive agent can appear to improve because search keeps a lucky component, even if its optimizer does not improve. We ask whether an optimizer revision is worth retaining: old and revised instructions learn from the same parent, their descendants freeze before a fresh gate opens, and both branches are charged. Across nine Skill source runs, gated TriEvo has a +2.53-percentage-point mean difference from a duplicate-incumbent tournament (source-mean 95% interval, 1.83–3.23). After all Policy components are reset, retaining its gated optimizer state yields a +3.02-point difference from accept-all state (1.98–4.07); a GEPA-on-Evolver adapter has the same direction. The original Prompt/Skill/Memory panel has a +7.53-point macro difference, but Memory is negative. These endpoint comparisons motivate a prospective selection-and-reuse procedure. They do not establish superiority to the adapted GEPA author, universal transfer, or a cost advantage. Task-population generalization remains outside the reported aggregate analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.