acceptodds
Under review as a conference paper at ICLR 2027

What to Learn, What to Keep: Information-Guided Sampling and Text-Space Weight Decay for Self-Evolving Agent Skills

Abstract

Large Language Model (LLM) agents fail on complex, domain-specific opera- tions that require heuristics absent from pretraining data. Skill-file self-evolution addresses this by iteratively rewriting a skill file from failure analysis, but exist- ing methods sample training instances uniformly at random, wasting budget on redundant or uninformative failures. We propose Information-Guided Active Sampling (IGAS), which casts instance selection as active learning over a non- parametric (text) skill: a multiplicative acquisition utility combining a failure- signal factor (failure similarity shaped by an entropy peak) with task represen- tativeness, and utility-weighted facility-location batch selection under a mono- tone submodular coverage objective with a (1 − 1/e) greedy guarantee. Be- cause a purely informative sampler drifts toward atypical instances whose im- provements do not transfer, we introduce ε-anchored mixing and a re-verification quota. Our evaluation protocol uses a four-way split (Train / gating ValA / final-selection ValB / Test) with strict gated acceptance (∆ > 0), a cross- round rejection memory, and a single post-hoc Test evaluation; it also exposes an inflation-accumulation phenomenon in single-validation-set selection, which our Best-of-N final selection counters through same-window dual evaluation and a paired-bootstrap discount. We evaluate on six datasets (LiveMathematician- Bench, SearchQA, DocVQA, SpreadsheetBench, ALFWorld, OfficeQA) under a preregistered controlled two-arm design—the original SkillOpt baseline ver- sus our full method—with 10 seeds per dataset per arm (120 runs in total).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.