acceptodds
Under review as a conference paper at ICLR 2027

Policy-Invisible Violations in LLM-Based Agents

Abstract

We identify policy-invisible violations: an LLM agent carries out a user's request as intended, yet violates internal organizational policies. The action appears appropriate from the agent's visible context, but whether it is permitted depends on organizational facts outside that context—such as recipient access rights, document-sharing restrictions, or relevant session history. These failures arise during ordinary, cooperative use, without adversarial manipulation. We introduce PhantomPolicy, a diagnostic benchmark comprising 494 enterprise-policy test cases across eight categories. In 1,080 paired-world episodes, observed legitimate completion is 64.2% with hidden state, 91.0% with queryable state and 95.0% with visible state; unsafe proposals are 64.8%, 7.2% and 8.9%, respectively. Caution under hidden state lowers both completion and violations, while model-level results show that access alone does not ensure correct use of policy facts. Two follow-up controlled studies support the benefit of relevant state. Our reference solution, Sentinel, checks proposed effects against policy and state and jointly plans evidence acquisition and goal-preserving repairs. In controlled continuing workflows, observed unsafe execution is 19.4 percentage points lower with Sentinel, and conservative missing-outcome bounds preserve the reduction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.