Execution-Grounded Credit Assignment for Self-Evolving Agents
Abstract
# Abstract Recent research on the self-improvement of language model agents has shown that iterative refinement guided by execution feedback is effective. On long-horizon tasks, however, credit assignment remains difficult because task outcomes reflect many interdependent actions. After a change, an agent may succeed without exhibiting the behavior the change was intended to produce, or fail despite exhibiting it. We propose an evolution method that uses execution-grounded component-level credit assignment to guide changes to instructions, tool code, and skills. Context-Anchored Gene Mutation derives mutation intents from outcome-associated behavioral patterns in the parent agent's execution histories and combines them into plans constrained by evolution context accumulated over earlier rounds. Applying these plans to the parent generates candidate agents, with each change linked to its intent. To explore diverse mutation candidates while limiting costly full task evaluations, we use simulation-based candidate pruning. It simulates the parent and each candidate for a bounded number of steps from recorded decision points relevant to each intent, under matched language-model request conditions. Candidates that express the intended behavior more strongly than the parent in these simulations are prioritized for full task evaluation. After evaluation, Gene Distillation combines each change's source pattern, intent, observed behavior, and task outcome to update its credit. Stronger evidence of expression increases the weight of both positive and negative outcomes. Credit accumulates in the evolution context to guide subsequent mutations. We evaluate the method on SEAL-0, which requires iterative retrieval and verification of potentially conflicting web evidence, and on EnterpriseOps-Gym, which tests stateful planning and tool use in enterprise workflows with interdependent steps. After evolution on five SEAL-0 tasks, accuracy on all 111 tasks, including the development tasks, rises from 41.44% to 63.36%; GEPA (adapted) reaches 50.45% on the same 111 tasks. After evolution on eight EnterpriseOps-Gym tasks, the evolved agent is evaluated on 642 tasks, achieving a task pass rate of 35.67% and a verifier-level pass rate of 68.9%, compared with 26.95% and 59.3%, respectively, for the initial agent. These results suggest that credit assignment for long-horizon agent evolution can be grounded in observed behavior as well as task outcomes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.