acceptodds
Under review as a conference paper at ICLR 2027

When Is a Policy Violation a Model Failure? Policy Observability and Failure Attribution in Agent-Safety Evaluation

Abstract

Reading a policy violation as a failure of rule following assumes that the agent had the rules and facts it needed. We study policy observability: whether the information available at decision time determines the authorization label used for evaluation. We formalize two conditions (sufficiency, that the scored label is a function of the agent-visible input, and presentation invariance, that this input does not change under transformations the benchmark’s own data format declares meaningless) and give an audit protocol with checkable certificates. An audit of 7,974 released AuthorityBench records resolves 65 call-level replay discrepancies through a documented reconstruction of historical scenario definitions. We also construct 34 policy pairs across 21 scenarios with identical agent-visible inputs under the released input specification but opposite authorization labels, so the available information does not determine the required decision. Applied to seven further released agent benchmarks, the protocol finds sufficiency counterexamples in two, presentation that varies with process state in the released pipelines of three, and one benchmark that scores infrastructure errors as successful attacks. In one of these pipelines, changing only the interpreter’s hash seed changes a deterministic agent’s scored utility on 6 of 35 affected cases (all in one user task). Experiments on an LLM authorization judge complement the audit. In 330 evaluation attempts on 55 action proposals from three retail tasks, candidate-first ordering had zero estimated excess disagreement relative to generic canonicalization in 50 complete-valid comparisons. The disagreement that did occur, on three proposals where the policy text and the tool description support different decisions, replicates in a pre-registered follow-up: reordering the top-level members of content-identical evidence flips the evaluator’s decision in every fresh session, 52 random member orders of one proposal split 30 to 22, and randomizing only nested members never changes it (0 of 12, against 7 of 12 for top-level orders). When the conflict is removed from the text (one of two pre-registered variants), the split disappears (51 of 51 orders approve). This evaluator fails our competence control (14 of 20, partly on flawed items); the one evaluator that passes it, Claude Opus 5.5, approves on all 52 random orders. These experiments measure decision consistency, not authorization accuracy, on proposals selected because they flipped. Together, the findings identify information, versioning and measurement conditions that must be checked before attributing benchmark violations to model failure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.