acceptodds
Under review as a conference paper at ICLR 2027

Skill Libraries Change What an Edit Does: When Should Agent Skill Validation Run In Context?

Abstract

Self-improving language agents keep a skill library of learned instructions, and they keep editing it. Every edit must pass one small test before it enters the library: a validator runs the candidate on a few episodes and accepts it only if performance improves. Current systems run this test in isolation, although at deployment the edit will live inside the library, surrounded by other instructions that share its prompt. In principle that difference matters: prompt context is known to change model behavior. In practice, on ScienceWorld with a 7B model, we present the first controlled measurement of that difference. Library context does reshape edit effects, shifting an edit's measured benefit substantially across contexts, yet in none of the 24 controlled units does it change the accept/reject decision at a practical threshold. A simple risk analysis explains both facts: a verdict survives a context change unless the distortion of isolated testing, or the noise of a small test, rivals the edit's true effect. Pre-registered equal-cost experiments test this analysis; the verdict is budget-dependent. At small budgets both validators are unreliable: run-to-run noise, not context distortion, is the main bottleneck, and both frequently misjudge the sign of an edit's effect. At full budget neither shows an advantage. On 64 edits produced by real updaters the two are statistically equivalent at every decision threshold a deployment would act on, within the single-model, score-granular scope tested here. The result is a simple routing rule: validate in the deployment context only when the distortion of testing in isolation exceeds the noise of a small test; otherwise isolated validation is the cheaper correct choice.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.