Paired-Contract Audits for Temporally Valid Tool-Agent Leaderboards
Abstract
Tool-agent leaderboards guide system selection, but a score measured against one interface snapshot need not preserve the decision after the interface changes. We introduce PACTBench (Paired Audit of Contract Transitions), an audit protocol that turns this temporal-validity question into a controlled comparison. PACTBench pairs contracts that preserve a task's goal, initial state, and validator; commits the later contract before execution; and reevaluates a predeclared policy panel under a shared observation boundary, backbone, and budget. Its rank-transition signature records score-band width, pairwise reversals, tie-aware rank agreement, and the winning set. The protocol makes the information boundary explicit: ordinary policies may interact with the revised tool only through the declared runtime channel, while the updated static contract and validator remain sealed. We give a task-blocked estimator, a reporting contract for external execution surfaces, and an audit record that separates public contract evidence from execution evidence. This design makes temporal rank preservation a concrete requirement for benchmark-supported tool-agent selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.