acceptodds
Under review as a conference paper at ICLR 2027

EvoClawBench: Can Agents Learn Reusable Skills from Their Own Runs?

Abstract

Existing agent benchmarks primarily test task completion, tool use, or skill utility, but do not isolate whether a runtime can convert evidence from its own runs into reusable skills that improve fresh executions after authoring overhead. We introduce EvoClawBench, a benchmark for this closed-loop skill-learning question on repeated, fixture-backed tasks. EvoClawBench compares direct execution without skills, PreSkill authoring before execution, and PostSkill summarization from first-run evidence followed by a fresh second execution of the same task. The suite contains 100 tasks and 502 sub-problems across coding, data, office, security, operations, and domain-document workflows, with support for multiple agent runtimes. We evaluate five models under OpenClaw and nanobot on the full 100-task suite, report percentage-point changes against a skill-free rerun inside each run, and log execution status for timeouts, empty transcripts, and fallback replies. Self-authored skills change mean scores by at most about 3.7 points on the four stable model rows and by about 5–6 points on Qwen3.6-Plus, comparable to the -3.3 to +3.1-point gap between two skill-free executions of the same model; DeepSeek-V4-Pro is the exception, with nanobot skill-mode scores near 19% while skill-free reruns stay near 82%. Reuse is not cheaper: excluding the first run, the summary and second execution still use about 2.6× the tokens of a direct run. These results indicate that learning reusable skills from an agent's own runs is selective and cost-sensitive, rather than an automatic benefit of adding skill authoring to an agent loop. We release the code at https://anonymous.4open.science/r/EvoClawBench-9380/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.