ProSEvo: Self-Evolving Agents Without a Stronger Teacher
Abstract
Can an agent improve from its own interactions without a stronger model to teach it? We study this question with a single locally hosted generative backbone used for experience collection, offline compilation, and subsequent agent generations. The agent must select self-generated experience carefully: a successful trajectory can still contain erroneous steps. ProSEvo uses environment outcomes, source-linked event labels, and deterministic checks to select what persists. Supported procedures become typed memory for the role that will use them; admitted original outputs become memory-free supervision for the role that produced them. Each updated student then collects experience for its successor, so both memory and parameters evolve across rounds. Under a unified same-backbone protocol on OfficeBench, AppWorld, and -bench Retail, ProSEvo raises macro task success from % to %, versus % for the strongest controlled comparator, % with memory alone, and % with parameter updates alone. It also leads the compared methods on each benchmark under this protocol. Fixed-checkpoint and round-wise controls distinguish the contributions of compiled memory and cumulative student-generated experience.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.