PROSE: Compact and Traceable Skill Evolution
Abstract
LLM agents increasingly use skill documents to reuse experience across related tasks, but automatically maintaining these documents remains challenging. A fully autonomous loop must identify lessons that apply beyond a single task, incorporate them without uncontrolled growth, and retain an inspectable basis for each revision. We introduce PROSE (PROcedural Skill Evolution), an experience-driven method for training the skill document of a frozen agent. Inspired by group-relative optimization, PROSE contrasts repeated attempts of the same task for task-conditioned attribution, then uses a batch-level optimizer to consolidate recurring mechanisms across tasks. It turns these lessons into localized updates to a structured skill document, with programmatic evidence and budget checks that bound document growth and keep each change traceable to recorded experience. Starting from an empty document, PROSE improves held-out success by to points across SpreadsheetBench, AppWorld, and LiveMathematicianBench. It attains the highest mean on every reported metric among the evaluated skill-learning methods. It also produces the most context-efficient skill documents, 43–88% smaller than those of competing methods. On AppWorld, this also holds when a single open-weight 27B model fills every role. We will release our implementation as open-source code. These results support skill documents as a compact and traceable adaptation layer for frozen agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.