acceptodds
Under review as a conference paper at ICLR 2027

Eviskill: Grounding Skill Evolution in Replayable Evidence

Abstract

Continual skill evolution enables LLM agents to accumulate and refine reusable procedural knowledge from interaction experience without updating model parameters. Recent experience-driven methods use execution trajectories to create and iteratively revise such skills. However, edits generated from execution trajectories do not consistently produce their intended behavioral improvements, and revisions rejected by global validation may still contain useful constituent edits. In this paper, we introduce Eviskill an evidence-driven framework that organizes continual skill evolution into evidence construction, behavioral verification, and cross-epoch refinement. Specifically, Eviskill first transforms localized execution observations into Replayable Evidence Cards and synthesizes candidate edits with explicit links to their supporting task contexts and trajectory ranges. It then re-executes the referenced behaviors under the edited skill to verify and refine these edits. Finally, global validation determines whether the resulting revision updates the Validated Skill, while replay-supported edits from rejected revisions are retained provisionally and unresolved evidence is carried across epochs for further refinement. Experiments on three interactive benchmarks across six LLM backbones demonstrate consistent improvements over skill evolution baselines. Our code, datasets and implementation details are available at https://anonymous.4open.science/r/eviskill-22B8.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.