Autonomous Strategy Exploration for Self-Evolving LLM Agents
Abstract
LLM agents have achieved impressive performance on complex tasks, including interactive environments involving long-horizon decision making. Nevertheless, most agents are unable to continually improve their behavior from experience at test time. Self-evolving agents offer a promising alternative by retaining memories and reflections across episodes, enabling behavioral adaptation without modifying model parameters. A key challenge, however, is that accumulated experience can progressively narrow an agent's behavioral diversity: successful strategies are repeatedly reused, while potentially better but unexplored alternatives become increasingly unlikely to be attempted. We refer to this phenomenon as exploration collapse. To mitigate this issue, we introduce an approach that explicitly organizes the agent's strategy space as a directed acyclic graph, where nodes represent milestones and edges capture prerequisite relationships. The strategy space is continually expanded by identifying evidence-supported yet unexplored directions, while a dedicated selection mechanism balances pursuing promising known strategies with exploring new ones during planning. The proposed approach consistently outperforms all evaluated baselines. Extensive ablation studies further confirm the importance of its individual components and demonstrate robust performance across diverse settings, showing that explicitly maintaining and exploring a structured strategy space can support sustained exploration in self-evolving agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.