Same Requirements, Different Decisions: How Presentation Changes Which Requests LLM Agents Allow
Abstract
LLM agents increasingly decide whether an action is allowed under an organization's policies. Policies list their requirements, yet they rarely say how these requirements combine into one decision, for example whether every applicable document must permit the action before it proceeds. The agent must therefore infer this unstated rule. We examine whether LLMs reach the same decision under every presentation of the same requirements. We compare the same requirements under six presentations that leave every correct decision unchanged, from layout and wording to the number of documents and the output requested. Across 12 open-weight and three Claude model configurations, models do not apply one rule. Layout, wording, and length move the rate of wrongly allowed requests by at most 5.4 percentage points, in both directions. By contrast, splitting one policy into two documents, or asking for supporting fields with the decision, raises that rate by 22 to 32 points, almost entirely toward wrong approvals. The errors consistently reflect the rule each presentation suggests. Reasoning models, for example, treat the two documents as alternatives. Some models require every condition in a bare list to hold and refuse requests they should allow. Claude Opus 5 makes no errors on plain-text policies. Yet Opus 5 wrongly approves a request in 32 of 72 attempts once one organization's rules are split across two of its own documents and supporting fields are requested. The same split with only the decision requested produces no wrong approval. The presentation thus supplies the decision rule that the policy leaves unstated. A test of a policy-following agent under one presentation therefore does not show how the agent will decide under another.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.