EvoGap: Weakness-Graph-Guided Skill Evolution for Frozen LLM Agents
Abstract
Frozen agents can improve through external skills, but what they learn is constrained by the tasks they encounter. Generating additional tasks can broaden practice, but may not sufficiently target diagnosed behavioral weaknesses. Consequently, agents may accumulate more experience without systematically repairing the weaknesses that actually limit their performance. To overcome this limitation, we propose EvoGap, a weakness-graph-guided framework that turns diagnosed failures into executable practice and uses validated outcomes to guide skill evolution and subsequent practice. An updatable weakness graph records behavioral gaps and repair evidence to select the next target. Depending on a weakness's repair state, EvoGap runs breadth-oriented or depth-oriented practice to generate executable tasks with verifiable outcomes. Environment feedback informs the extraction of environment knowledge and behavioral skills. Candidate skills are admitted only after targeted validation, and their measured effects are written back to the same graph. This feedback loop connects practice selection, skill updates, and the next round of practice. Across AppWorld, BFCL-v3, and the three -Bench domains, EvoGap consistently improves task success over prior memory- and skill-based baselines with a frozen backbone, by points on average and up to points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.