Policy-Invisible Violations in LLM-Based Agents
Abstract
We identify policy-invisible violations: an LLM agent carries out a user's request as intended, yet violates internal organizational policies. The action appears appropriate from the agent's visible context, but whether it is permitted depends on organizational facts outside that context—such as recipient access rights, document-sharing restrictions, or relevant session history. These failures arise during ordinary, cooperative use, without adversarial manipulation. We introduce PhantomPolicy, a diagnostic benchmark comprising 494 enterprise-policy test cases across eight categories. In 1,080 paired-world episodes, observed legitimate completion is 64.2% with hidden state, 91.0% with queryable state and 95.0% with visible state; unsafe proposals are 64.8%, 7.2% and 8.9%, respectively. Caution under hidden state lowers both completion and violations, while model-level results show that access alone does not ensure correct use of policy facts. Two follow-up controlled studies support the benefit of relevant state. Our reference solution, Sentinel, checks proposed effects against policy and state and jointly plans evidence acquisition and goal-preserving repairs. In controlled continuing workflows, observed unsafe execution is 19.4 percentage points lower with Sentinel, and conservative missing-outcome bounds preserve the reduction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.