Measuring the Marginal Utility of Skills at the Execution State
Abstract
Skill libraries package procedural knowledge for agents, but retrieval ranks relevance rather than testing execution. We define the marginal utility of a fixed package-plus-announcement protocol at execution state as its change in expected verified reward. At each checkpoint we replay one realised prefix, apply a protocol arm, independently continue it, and compare verifier rewards. Across 22 screened realised states from eleven SkillsBench tasks under one model and agent (516 runs), the health-gated own-package contrast was +0.147. Conditional on these fixed states, a runs-only 95% interval was [+0.082, +0.213] and within-state randomization gave ; an 11-cluster interval crossed zero [-0.020, +0.315] (), making cross-task mean inference method-sensitive. Utility did not separate checkpoint rules, but uptake differed by +39.9 points between first-write-rule and three-fifths checkpoints; two no-write fallbacks and one reversed pair limit that label. Only two tasks had a foreign top-1. The contribution is a signed, state-conditioned protocol and auditable finite-sample map, not a router.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.