PRA: Certifying Skills Before They Become Tool-Agent Actions
Abstract
Tool agents repeat short action sequences, and every step of such a sequence costs a policy call. Recurring behaviour, which we call a skill, is therefore an opportunity to reduce calls. A large body of work asks what is worth constructing from demonstrations. However, whether a deployed policy should then launch that skill with one decision is a different question. Existing criteria cannot answer it, because recurrence and executability describe the demonstrator and the environment, not the policy that will pay. On ALFWorld, catalogues selected for recurrence or executability realise their declared effect more than 90% of the time yet cut task success from 0.861 to as low as 0.493. To address this, we propose policy-relative admission (PRA). PRA separates skill construction from action promotion and settles promotion as a decision about one frozen policy. Candidates are lifted into operator sequences that a deterministic binder executes blind, so runnability is measured rather than assumed. Certification then asks two questions on held-out games. One asks whether committed execution realises the declared effect, and the other whether exposing the action lowers decisions without losing task success beyond a non-inferiority margin. Experimental results show that the same behaviour is an action for one policy and not for another, since the admitted set moves with policy capacity and family and with how the policy was trained. Both certificates reject distinct failure modes, and neither subsumes the other. Under a behaviour-cloned policy and a released RL-trained agent alike, PRA is the only method that cuts deployment-time decisions at non-inferior success in both environments. On WebShop it removes 54% of the cloned policy’s decisions and raises its official score from 0.248 to 0.395. Our work turns promotion from a property assumed at construction time into a decision a deployment can audit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.