RRU: Improving Agent Skill Generation with Reflect-Replay-Update
Abstract
Agent skills extend language agents with workflows and knowledge beyond their pre-trained capabilities. Existing skill generation frameworks adopt a Reflect- then-Update approach: the agent solves a batch of tasks, reflects on the rollouts, and updates the skill to capture recurring patterns and resolve conflicts. This ap- proach can let rollouts from a single batch remove an instruction supported by earlier batches (overwriting) or become an instruction applied to all tasks (over- generalization). These failures arise because each update conditions only on the current batch, as earlier rollouts are often discarded after reflection. We introduce Reflect–Replay–Update (RRU), which improves grounding by maintaining com- pact records of the rollouts that provide evidence supporting each instruction and replaying these rollouts during skill updates. With replay, an update can revise an earlier instruction rather than remove it, and generalize a new instruction only when earlier batches support it. On AppWorld, SpreadsheetBench, and Legal Agent Bench with two open-weight models, RRU outperforms strong baselines (Trace2Skill, SkillOpt, and Combee) by an average of 8.5% on AppWorld and 5.6% on SpreadsheetBench across the two models, and by 3.7% and 1.2% on Le- gal Agent Bench with Qwen and Gemma. Ablations show that replay improves accuracy by 5.0% on average and reduces variation across seeds in three of four settings, and that compact records reach higher accuracy than complete rollouts in three of four settings and use fewer tokens in three of four settings (up to 30%).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.