What Agent Skills Add: A Causal Audit of Task Conventions
Abstract
Does a Skill improve an agent's capability, or supply a task convention that the evaluation expects but the task omits? We audit this distinction through matched content interventions and prospectively frozen claim gates. Across 378 retained rollouts, a benchmark-native pilot nominates an unstated projection rule; a frozen 54-rollout removal/insertion experiment finds 36/36 hidden distance-clause passes with exact guidance versus 0/18 with generic controls. A 270-rollout documentation-by-Skill factorial in five constructed families finds a 0.467 correct-minus-generic gain when documentation omits the rule (simultaneous interval [0.300, 0.633]; adjusted p = 0.000778), but only three families have positive increments, missing the registered four-family recurrence gate. Harm and repair gates also fail. An independent 1,152-opportunity operational test leaves information-source attenuation unresolved (0.085938, 95% interval [-0.208456, 0.380331]); each condition has 97/128 zero-scored process failures. A fresh external-rule factorial across eight source lineages and 1,536 opportunities also misses its recurrence gate: source-information difference 0.053711 (simultaneous interval [-0.097362, 0.204784]). Its matched Skill-minus-document difference is 0.001628 ([-0.149445, 0.152701]), within a prespecified broad ±0.25 material-equivalence margin, not proof of equal carriers. The evidence supports a local effect of convention-and-sequence guidance and a finite-bank gain, not a general mechanism or authoring rule. Skill evaluation must distinguish supplying missing task information from recurring or carrier-specific value.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.