acceptodds
Under review as a conference paper at ICLR 2027

TriEvo: Three-Agent Evolution of Policies and Optimizers

Abstract

An adaptive agent can appear to improve because search keeps a lucky component, even if its optimizer does not improve. We ask whether an optimizer revision is worth retaining: old and revised instructions learn from the same parent, their descendants freeze before a fresh gate opens, and both branches are charged. Across nine Skill source runs, gated TriEvo has a +2.53-percentage-point mean difference from a duplicate-incumbent tournament (source-mean 95% interval, 1.83–3.23). After all Policy components are reset, retaining its gated optimizer state yields a +3.02-point difference from accept-all state (1.98–4.07); a GEPA-on-Evolver adapter has the same direction. The original Prompt/Skill/Memory panel has a +7.53-point macro difference, but Memory is negative. These endpoint comparisons motivate a prospective selection-and-reuse procedure. They do not establish superiority to the adapted GEPA author, universal transfer, or a cost advantage. Task-population generalization remains outside the reported aggregate analysis.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.