PACT: Provenance-Aware Contract Testing — What Do LLM-Agent Task-Contract Checkers Actually Certify?
Abstract
LLM agents now act on behalf of users: sending messages, initiating payments, and booking travel. To keep them in check, developers run contract checkers: programs that verify the agent only acted on destinations named in the task text. A clean certificate is usually taken to mean the agent did only what the task authorized. But the checker only matches strings: it can pass an action the task never authorized, and flag one the task did. We measure that gap on 7,598 public AgentDojo episodes. We first compute exact risk certificates for such a checker. We then test what its flags mean: independent LLM judges blindly review 400 stratified trace-task pairs and decide whether each sampled flag is a truly unauthorized action, under a predeclared rubric. These two steps — certify the label, then measure its meaning — form an audit framework we call Provenance-Aware Contract Testing (PACT). On the attack side, the checker is largely sound: about three-quarters of sampled attack-side flags hold up as genuinely unauthorized actions, and a predeclared extension covering what the main audit cannot see shows that almost no attacks are missed. Yet the dominant failure is over-strictness: 57.5% of sampled benign-side flags are actions the task did authorize (the exact rate depends on the coding rubric). A provenance-aware resolver, which tracks where each destination string came from rather than whether it appeared anywhere, authorizes 2.6x fewer semantic violations than the breadth baselines and eliminates all measured false rejections. The frame-level semantic violation rate is estimated under measured worst-case bounds, not certified. The rubric is human-validated on every checkable boundary unit, with zero residual disagreement, and in a prospective test on unseen tasks, the compiled resolver falls short of its predeclared threshold. Certificates should be read as statements about what the checker actually verified, not as proof of what the task authorized; we measure exactly what separates the two.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.