acceptodds
Under review as a conference paper at ICLR 2027

Quantitative Certification of Agentic Tool Selection

Abstract

Large language models (LLMs) are increasingly deployed in agentic systems, where a fundamental task is mapping user intents to relevant external tools. Errors in tool selection can have severe outcomes, such as unauthorized data access, even without modifying the agent's underlying model. Existing evaluations measure performance on curated, benign benchmarks. However, a pipeline's behavior in deployment depends on the tool pool the agent actually encounters, which in open registries is shaped by third parties. We introduce CATS (Certification of Agentic Tool Selection), the first statistical framework that returns high-confidence upper bounds on the probability that a tool selection pipeline satisfies a declared safety specification under a plausible tool distribution. CATS reduces the certification of tool selection to a Bernoulli estimation problem, drawing sequences of inserted tools from a declared distribution paired with the safety specification. To model realistic deployment conditions, we instantiate this distribution as a stochastic process that generates sequences of inserted tools round by round, conditioning each round on the agent's selection in the previous round. CATS aggregates the trial outcomes into a one-sided Clopper-Pearson upper bound on the probability that the specification is satisfied. By returning this bound as a certificate that holds under the distribution of inserted tools, CATS makes safety claims intuitive, actionable, and comparable across models, retrievers, mitigations, and registry policies. Across popular BFCL and OpenAPI tool pools, CATS shows that current LLM agents remain fragile under Distractor Selection and Top- Saturation specifications: their certified upper bounds drop to approximately 20%, far below their lower bounds on the clean pool.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.