Continual-Skill: Utility Optimization-Driven Co-Evolution of Policies and Skill Libraries
Abstract
Self-improving agents continually accumulate reusable skills, yet they lack an efficient mechanism for determining whether a stored skill still benefits the current policy. Semantic relevance does not reveal a skill’s actual effect on task performance, while explicit with/without-skill evaluation requires costly additional environment interactions. We introduce **Continual-Skill**, a framework that turns ordinary training trajectories into online evidence of skill utility. Its key idea is to randomize skill exposure during training, creating exposed and unexposed observations for candidate skills, and to use inverse-probability weighting to estimate their marginal effects on agent performance. The resulting utility estimates are updated alongside the policy and directly govern skill routing and lifecycle decisions, including promotion, suspension, and archival; persistent task-domain failures additionally trigger new-skill generation. This design allows policy learning and skill reassessment to share the same interaction data, avoiding separate per-skill evaluation rollouts. On ALFWorld and WebShop, Continual-Skill remains robust under increasingly interfered skill banks and improves success rate by up to 40.4 percentage points over strong skill-learning baselines on expanded WebShop banks. Results suggest that continual agent improvement requires treating skill utility as a policy-dependent quantity that must be repeatedly identified and updated throughout learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.