Accepted but Not Enforced: Auditing Tool-Call Policy in LLM Agent Frameworks
Abstract
Agent frameworks expose policy loaders, guardrails and approval hooks to control tool execution, but accepting a policy does not ensure that it governs the resulting effects. We study this gap through a proposal–decision–effect authorization chain that records policy activation, authorization decisions and committed effects as separate observations, and we state the enforcement contract a deployment must meet before committing a protected effect. Auditing four pinned deployments, a five-framework route study and an eighteen-framework census, we find that contract broken at every position it names: in two native policy languages, 10 of 13 clause instances that loaded without governing a decision produced no diagnostic, and an alternative tool path escaped the configured check in four of five frameworks, the fifth covering both tested paths with one kernel-level filter. We then name the primitive each obligation needs and build them natively in two versions of OpenAI Agents SDK. Waiting reservations preserve feasible legal effects that immediate refusal loses, and explicitly propagated child hooks enforce policy on nested calls. Across 150 runs per configuration per version, the integrated layer takes violating runs from 75 to zero and runs that lose feasible legal effects from 15 to zero, restoring enforcement without blocking the work a deployer wants done.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.