acceptodds
Under review as a conference paper at ICLR 2027

When Can a Skill Library Trust Its Execution Records?

Abstract

A skill library may use a successful execution record to predict an update's effect, decide whether to deploy it, or justify replacing the full skill with shorter text. These are different judgments. On sixteen known QA questions, records of the same proposal at other receivers reduce squared F1-effect risk by relative to equal-budget records of other proposals at the same receivers. That information raises the value of a fixed retrospective gate by F1 points over its matched control, yet the gate's own value relative to rejecting every update is estimated at points ( interval ). Receiver-only and pooled histories also attain lower absolute prediction risk than same-proposal history. In a separate crossed study, three local trials beat pooled records in two settings and lose in two others as the balance of local execution noise and nonlocal error reverses. Finally, performance substitution—replacing a full skill with one section while maintaining its mean score within a prespecified margin—changes answer distributions and source outcomes under GLM. Execution records do transfer useful information, but prediction, deployment, and retention of particular successes require distinct comparisons.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.