SkillGRACE: Group-Relative Advantage-Guided Contrastive Editing of Agent Skills
Abstract
Large language models (LLMs) have evolved from text generators into agents that reason, use tools, and interact with environments, while skills improve their behavior through reusable instructions without updating model parameters. In this paper, we argue that existing skill optimizers do not fully exploit the information in task trajectories: each execution of a validation task provides only a single reward signal to guide patch selection, and their computational cost increases substantially as task execution becomes more complex. Moreover, aggregating evidence across diverse tasks can obscure the behaviors responsible for success or failure, making the proposal of effective candidate edits challenging and thereby highly dependent on strong external LLM editors. To address these limitations, we introduce SkillGRACE, which formulates skill editing as group-relative policy optimization over an external language policy. To enable discrete skill optimization, SkillGRACE samples multiple rollouts for each task and leverages group-relative advantages to construct same-task contrastive evidence for patch generation. The proposed patches are then selected by replaying these trajectories under candidate skills, eliminating the need for a separate validation set and substantially reducing optimization cost. Notably, since each skill patch is informed by contrastive evidence from same-task rollouts, SkillGRACE enables the executor itself to generate effective edits without requiring a powerful external LLM editor. Across three benchmarks, SkillGRACE consistently outperforms the strongest baseline in each setting. Across three benchmarks, SkillGrace outperforms SkillOpt by 2.8 percentage points on average and by up to 6.8 points, while using 77.6% fewer optimization rollouts overall. An ALFWorld ablation further shows that same-task contrastive evidence improves success by 13.4 points over cross-task collections, highlighting its importance for effective self-evolution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.