acceptodds
Under review as a conference paper at ICLR 2027

Skill Lock-In: Silent Retirement Failures in Self-Evolving LLM Skill Libraries

Abstract

Self-evolving LLM agents can retain harmful advice in growing skill libraries despite continued evaluation. Judge bias affects skill removal and admission differently. When feedback weakens measured effects without changing their sign, zero remains a meaningful decision boundary, while a permissive admission rule can allow much greater true harm. We identify sparse skill use, optimistic feedback and routing changes as interacting causes of this skill lock-in. We introduce PAWL (Paired Anytime-valid Watchdog for Libraries), a lifecycle that estimates a skill's contribution on the tasks where it is used. It randomly withholds skills to maintain task-matched no-skill references and uses confidence bounds valid under repeated inspection to govern removal and admission. Experiments across four solver families and two live LLM judges show stronger harmful-skill detection and more selective admission. In one matched optimistic-feedback setting, PAWL retires 16 of 18 planted harmful skills, versus two for the success-count baseline. Separate replays show that normalizing success counts by skill use remains insufficient. On screened, loop-synthesized coding skills, PAWL admits 15 of 18 helpful candidates and none of 18 harmful candidates under strong simulated optimistic feedback within the tested budget. These findings establish measurement, threshold choice and sustained exposure as central design considerations for self-evolving skill libraries.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.