Don’t Forget What You Just Learned: Self-Directed Repair Makes Agent Skills Persist
Abstract
Intelligent agents should learn from their mistakes so experience compounds into continuous improvement. Yet today’s agents recover from a mistake only temporarily: in-context reflections are easily forgotten across sessions and compete for limited context, while full parameter fine-tuning is expensive and can interfere with existing capabilities. We introduce Self-Direction, a mechanism for agents to learn from mistakes that combines the benefits of temporary retries in-context and policy updates. When an agent fails a task, it reflects on its own trajectory in order to extract a lesson and immediately distills it into a lightweight low-rank adapter. By keeping the original weights frozen, Self-Direction preserves existing capabilities while only training \ 0.02% additional parameters for each adapter. The agent then saves these adapters into a persistent skill library, automatically routing them to solve related problems later. Experiments with open agents based on the Qwen-family across open-ended research (GAIA), tool-use workflows (Claw-Eval), and terminal tasks (Terminal-Bench 2.1) show our method improves overall success rates by 8.9 to 17.0 percentage points and performs competitively with frontier systems. We find Self-Direction avoids repeating mistakes across future encounters of a task, nearly doubling the reliability of an agent’s corrections and enables transfer of learned skills with just a handful of adapters generalizing to dozens of new tasks without additional training. Ultimately, this self-directed repair transforms everyday mistakes into lasting, reusable capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.