ToolTrials: Agents Learn Opaque Tool Contracts by Experiment
Abstract
An agent can read a tool's schema and still choose the wrong action: labels such as standard, compat, and strict often name distinct operational contracts whose semantics are visible only through execution. Existing approaches rewrite documentation or accumulate task trajectories, but the examples they observe need not distinguish competing behaviors. We introduce ToolTrials, which treats a new tool as an active system-identification problem. Given a finite executable candidate library containing the target and a legal query pool, it maintains candidate contracts, selects the input-policy trial that maximally partitions them, executes the trial, and compiles the identified mapping into a reusable behavioral card. Across 48 opaque implementations, eight behavior families, 288 held-out tasks, and 10,272 planned agent decisions, actively selected receipts reach 84.0-93.1% success versus 53.1-54.5% for passive observations. Compiling those receipts into cards reaches 98.3-99.0%, the oracle range. A matched control applies the identical compiler to model-designed probes and reaches 97.6-100.0%, localizing the downstream gain to verified contract compilation while ToolTrials supplies deterministic, auditable probe selection. Three independent passive draws preserve 44.3-50.7-point card gains. A prospectively frozen non-randomized study identifies 18/18 installed Python, Node, Ruby, and GNU provider/mode contracts and predicts 90/90 held-out outputs; two random trials identify 52.4%. On a separately blinded natural-mode study, opaque case identifiers eliminate answer cues: cards solve 120/120 decisions across three model families, versus 105/120 from schema alone and 107/120 from fixed receipts. Deterministic transfer additionally verifies 480 GNU-wrapper actions and 756 two-tool plans. Two stored trials detect all 480 tested policy-mapping changes, while five-fold repetition preserves 98.5-98.6% identification under 10% output noise. ToolTrials turns tool exploration from incidental trial-and-error into a small, planned, transferable experiment whose evidence and conclusion are executable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.