acceptodds
Under review as a conference paper at ICLR 2027

How Much Reward Does Self-Evolution Need? Pairwise Validation Without Task-Reward Feedback

Abstract

Self-evolving agents rely on task rewards to guide revisions, yet outside benchmarks labeled examples may be scarce, expert feedback costly, and automated evaluators incomplete. Can agents improve their prompts without reward feedback at every evolution step? We introduce PairEvolve, which removes task scores and gold answers from mutation and uses a frozen LLM to compare the original and revised outputs for acceptance. Two variants address different levels of reward access. PairEvolve-Val retains validation-score parent selection when a labeled set is available to choose which prompt to revise next. PairEvolve-Elo replaces that feedback with Elo ratings—numerical ratings updated from pairwise outcomes—so both acceptance and parent selection can proceed without task rewards. Both retain labeled final selection in our evaluation. Across four prompt tasks and two agent models, Val and Elo improve mean test scores over no evolution in five and six of eight settings, respectively, retaining full-reward performance in some settings and trading performance for reduced reward access in others. Validator studies identify how comparison format, thinking, and the treatment of undecided comparisons affect this feedback. Integrations with two program-evolution engines extend the comparison interface to code while retaining execution checks and proxy scores. Together, these results show how pairwise feedback can support evolution with fewer reward-dependent decisions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.