acceptodds
Under review as a conference paper at ICLR 2027

SkillCred: Online Value Estimation for Evolving Agent Skills

Abstract

Agent skills encapsulate reusable strategies and procedural knowledge to improve agent task performance. Recent work has begun to evolve such skills continually from historical interaction trajectories. However, skill evolution need not yield monotonic improvements. Newly generated versions may encode erroneous strategies, and blindly replacing an existing version with the latest one can degrade agent performance. To address this problem, we present SkillCred, an online framework that maintains competing versions within each skill cluster and estimates their values from trajectories generated during normal task execution. For each skill version invoked in an episode, a local progress assessor evaluates the agent behavior observed across all of its invocation intervals and produces one ternary local progress signal. SkillCred uses these signals to update a Dirichlet posterior over each version's local progress probabilities and applies Thompson sampling to route among competing versions while accounting for estimated value and uncertainty. It further uses posterior-based lifecycle control to promote, retain, or archive challengers according to posterior evidence and evaluation budgets. Because all value estimates are learned from ongoing task trajectories, SkillCred requires neither additional evaluation rollouts nor a held-out validation set. Experiments with three executor models on ALFWorld and WebShop show that SkillCred consistently achieves the highest mean success rate among the compared methods. Averaged across the three executors, its relative success-rate gains over the strongest baseline are 14.5% on ALFWorld and 38.2% on WebShop. The code is available at https://anonymous.4open.science/r/SkillCred.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.