TRACE: Enforcing Tool-Use Policies Beyond Model Guardrails
Abstract
Tool-using agents turn model outputs into actions that read from or change external systems, then feed the results back to the model. Model-side guardrails can discourage unsafe tool calls, but applications need a separate policy governing which calls may execute. We present TRACE, a policy-enforcement engine at the application's tool-call boundary. Before a tool call executes, TRACE checks it against a deployer-defined policy and returns ALLOW, ASK, or DENY. It records each decision separately from what the application reports as executed. We evaluate TRACE on 200 SWE-bench Pro tasks, testing each task under a protective policy and an allow-everything policy while holding the rest of the setup fixed. When the agent is instructed to manipulate grading, TRACE's protective policy cuts successful reward-hacking attempts from 50 to 14. Under ordinary task instructions, the policy has a small observed cost to legitimate work: the agent successfully completes 126 tasks, compared with 129 under the allow-everything policy. TRACE also supports an LLM as the decider for tool calls. We demonstrate this capability by replaying 5,650 tool calls with Claude Sonnet 5 as the decider. It denies 89 of 91 calls labeled as carrying attack payloads, compared with 75 for a basic rule-based decider, while denying 4 of 4,537 calls from honest runs. TRACE can combine both deciders while preserving every deterministic policy denial.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.