Repeatable but Not Replicable: Evaluator State as a Hidden Facet in Self-Evolving Agents
Abstract
Self-evolving agents improve by keeping new versions of their prompts or code that score higher on test tasks. Exact reruns often check such comparisons, but evaluators also write incidental strings, such as run directories and identifiers, into model inputs, and a rerun repeats them, hiding their effect. We treat these strings, the evaluator state, as a measurement facet: sharing a state cancels only the changes two versions share. In an unmodified SQL-writing agent, a new equal-length identifier changes 48 of 240 first queries but only 4 grades. Decisions on 100 tasks hold within a study, but on 10 tasks two evaluations of the same pair disagree on a median of 13 percent of decisions, and sharing a state did not measurably help. When a GEPA prompt-optimization loop shows the model its run directory, a rerun repeats every step with a database clock fixed, but other directories keep different prompts. On 12 validation tasks, the added lines lower the bar a new version must beat, and 14 of 18 loops keep a version worse on held-out tasks. Without the lines, these losses persist but mostly vanish at 100 validation tasks. Evaluations should rerun comparisons under fresh evaluator states.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.