Supervising Agent Evolution with Prefix Potentials
Abstract
Agent evolution through continued interaction and parameter updates requires informative supervision. However, environmental feedback is often sparse and concentrated at terminal outcomes, offering limited guidance for intermediate decisions. Typically, existing approaches alleviate this sparsity via separately trained evaluators or task-specific reference structures to construct step-level supervision. Instead, we propose **A**gent **S**elf-supervision through **C**ompletion-potential **E**stimation and **N**eighboring-prefix **D**ifferencing (ASCEND), which derives this fine-grained supervision from assessments by the evolving agent. Concretely, the evolving agent estimates a prefix potential at each step through an additional scoring call, assessing task-completion prospects based on the task objective and the actions and observations available up to that step. Differences between adjacent prefix potentials yield step-level signals, which are normalized, accumulated from each step onward, and combined with terminal outcome feedback to guide parameter updates. Evaluations on long-horizon interactive tasks and multi-hop reasoning demonstrate that ASCEND outperforms representative evolution baselines, with improvements extending to additional benchmarks without further adaptation. Further analysis shows that prefix potentials correlate with final task outcomes and that differences between adjacent prefix potentials align with independently assessed task progress. Code is available at *Anonymous*.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.