SKT: Self-Evolving Agent Skill Library with Plan-Driven Test-time Self-Correction
Abstract
Long-horizon LLM agents must plan and act over many environment turns, yet adapting them by updating parameters is costly and often impossible behind an API. Inference-time learning from the agent's own trials is the standard alternative, but memory-based methods keep experience as episodic text that is never compiled into procedures, code-based skill libraries apply only where actions are code, and step-level verifiers deliver their verdicts through the prompt, a channel that a small frozen policy can attenuate or ignore. To address this, we propose SKT, which evolves a natural-language procedural skill library from the agent's own successes and failures, consolidates failures across trials into new skills, and seats the same frozen weights in a separate verifier context whose corrections are executed directly as environment actions. Under a matched evaluation protocol with a frozen Qwen2.5-7B-Instruct, SKT consistently outperforms inference-time baselines on held-out ALFWorld tasks, approaches an RL-fine-tuned reference without any weight update, and ports unchanged to multi-hop QA and, on a commercial API backbone, to WebShop.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.