Governing Agent Self-Evolution: From Strategy Learning to Cognitive Organization
Abstract
Improving agent capabilities typically relies on manually written strategies and iterative debugging. Research on recursive self-improvement (RSI) seeks to automate this process, enabling agents to modify themselves based on execution feedback and allowing the improved systems to participate in subsequent rounds of optimization. However, as strategies accumulate, content that is irrelevant to the current situation or conflicts with other strategy content may interfere with task-related judgments, a phenomenon we refer to as knowledge contamination. Growing strategy content also increases inference costs during execution. We propose GASE (Governing Agent Self-Evolution), a framework that integrates governance throughout strategy formation and evolution, enabling self-evolution to continually produce information that supports subsequent strategy optimization and cognitive migration. From the outset of learning, the framework organizes strategies into units that can evolve independently, each specifying how the agent should behave under particular conditions, and continually establishes links between strategies and execution experience. Through these links, accumulated experience can be reused in subsequent evolution to provide, at low cost, successful cases as references for strategy revision, as well as validation examples and training material for migrating local condition-evaluation tasks. Condition-evaluation tasks validated for independent execution are then assigned to specialized executors, forming detachable cognitive modules. These modules retain links to their source strategies and versions, allowing the system to revoke their assigned evaluation responsibilities when source conditions change and return the corresponding evaluations to the main LLM. When integrated with EvoSkill and SkillOpt, GASE improves the task performance of both methods on ALFWorld, OfficeQA, and SearchQA. On ALFWorld, the full framework improves task success rates over EvoSkill and SkillOpt alone by 7.96 and 7.46 percentage points, respectively, while reducing the average number of main LLM input tokens per task by 10.95% and 26.03%, respectively. Experiments across multiple models and component ablations on ALFWorld further show that revising individual strategy units and reusing associated experience help improve strategy learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.