acceptodds
Under review as a conference paper at ICLR 2027

RRU: Improving Agent Skill Generation with Reflect-Replay-Update

Abstract

Agent skills extend language agents with workflows and knowledge beyond their pre-trained capabilities. Existing skill generation frameworks adopt a Reflect- then-Update approach: the agent solves a batch of tasks, reflects on the rollouts, and updates the skill to capture recurring patterns and resolve conflicts. This ap- proach can let rollouts from a single batch remove an instruction supported by earlier batches (overwriting) or become an instruction applied to all tasks (over- generalization). These failures arise because each update conditions only on the current batch, as earlier rollouts are often discarded after reflection. We introduce Reflect–Replay–Update (RRU), which improves grounding by maintaining com- pact records of the rollouts that provide evidence supporting each instruction and replaying these rollouts during skill updates. With replay, an update can revise an earlier instruction rather than remove it, and generalize a new instruction only when earlier batches support it. On AppWorld, SpreadsheetBench, and Legal Agent Bench with two open-weight models, RRU outperforms strong baselines (Trace2Skill, SkillOpt, and Combee) by an average of 8.5% on AppWorld and 5.6% on SpreadsheetBench across the two models, and by 3.7% and 1.2% on Le- gal Agent Bench with Qwen and Gemma. Ablations show that replay improves accuracy by 5.0% on average and reduces variation across seeds in three of four settings, and that compact records reach higher accuracy than complete rollouts in three of four settings and use fewer tokens in three of four settings (up to 30%).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.