The Contract–Injection Mismatch: An Execution-Conditioned Audit of Skill Self-Evolution
Abstract
Self-evolving skill frameworks write a reusable skill document from an agent's own task experience, and are accepted by a single average success rate: measured under one benchmark, one base model, and one deployment configuration, the evolution counts as effective once the number goes up. Whether this score represents the skill—whether it holds outside the configuration that produced it, and what it conceals within—has never been examined, yet self-evolved skills are being certified, compared, and deployed on it. We therefore conduct an execution-conditioned audit of SkillOpt, a representative self-evolution system: we compare skill and no-skill conditions on the same task pool, repeat every task five times, and read every rule evolution wrote. The audit finds that the score is a property of the execution condition: +13.4 points on SearchQA become −6.2 on OfficeQA, flipping or splitting again when only the base model or the deployment harness changes; even within the original configuration, the score is the residual of 315 tasks improved against 116 harmed. Separating the two sides: the gains are stabilization—turning intermittently correct tasks into always-correct ones—while the costs land on tasks that never failed, where injection introduces error types that did not exist before. The text explains both sides: evolution writes operating rules with preconditions, not new knowledge, yet injection places every rule on every task. Because a rule's preconditions can only reveal themselves in execution, the audit's rule of use is: do not inject by default; retry once with the skill on failure—outperforming both no-skill and full injection on all three benchmarks. In simpler terms: rules written for some executions are deployed into all of them, and whether they help is only knowable after execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.