acceptodds
Under review as a conference paper at ICLR 2027

SKILLFLOW: Benchmarking Continual Learning through Self-Evolving Agent Skills

Abstract

What distinguishes an agent that merely completes tasks from one that improves by completing them? Existing benchmarks primarily evaluate task success across isolated episodes or assess the utilization of pre-supplied skills, leaving unaddressed whether agents can convert their own experience into reusable procedural knowledge. We introduce SkillFlow, a benchmark designed for continual learning via self-evolving agent skills. Across 20 workflow families, its 166 tasks share execution structures while varying domains, artifacts, and operational conditions. Within each family, an agent begins with an empty skill library and can create, refine, and reuse skills derived from execution trajectories and verifier feedback. Using a unified Terminus-2 evaluation harness, we evaluate vanilla execution against an evolving-library setting across ten models. This comparison demonstrates a pronounced capability gap: models exhibiting comparable vanilla performance can differ substantially in their capacity to benefit from skill evolution, and frequent skill access does not necessarily yield effective adaptation. Trajectory analysis indicates that successful skill transfer relies on portable procedures and objective-aligned verification, whereas overly specific instructions and self-confirming checks are associated with negative transfer. SkillFlow renders these learning dynamics observable and provides a controlled testbed to investigate how and when agent skills foster continual adaptation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.