acceptodds
Under review as a conference paper at ICLR 2027

FAILURE2SKILL: LEARNING RESIDUAL SKILLS FROM AGENT FAILURES FOR INTERACTIVE TOOL USE

Abstract

Large language model agents can often repair a single mistake when given feedback, yet they rarely turn that mistake into reusable operational knowledge. We study a focused setting for self-improving tool-use agents: after an agent equipped with a current skill library still fails, can the residual failure be compressed into a new skill that transfers to related tasks? We propose Failure2Skill, a failure-driven skill induction pipeline that diagnoses failed trajectories, clusters causal residuals, and writes compact, guarded skill cards for later retrieval. Unlike trace-combination baselines that summarize successful and failed runs together, Failure2Skill treats failures as targeted evidence about missing preconditions, API response shapes, exception paths, and state-update constraints. On a 54-task AppWorld residual- family evaluation using a DeepSeek Chat code agent, Failure2Skill solves 54/54 tasks, while the best combined success/failure skill baseline solves 37/54. Across 18 task families, Failure2Skill wins 8, ties 10, and loses 0; its average verifier pass fraction is 1.000 versus 0.843 for the best combined baseline. Ablations show that removing individual residual rules often collapses final success, including payment-card retry, relation filtering, receipt attachment, note-format preservation, and comment parsing. These results suggest that failed trajectories are not merely negative examples: when localized, they are high-density supervision for building reusable agent skills

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.