acceptodds
Under review as a conference paper at ICLR 2027

Who Is Allowed to Act? Benchmarking Decision Rights in Document-Grounded AI Agents

Abstract

An agent’s ability to execute an action does not establish its right to do so. As AI agents begin to act on information spread across policies, approvals, records, budgets, and other documents, they must determine whether the available information actually gives them sufficient grounds and authority to take a proposed action. We define decision-right reasoning as the task of determining whether a proposed action is permitted, not permitted, or cannot be resolved from the available evidence. We organize this decision around six factors, covering the sufficiency of the evidence, the validity of the relevant authority, the reversibility and criticality of the action, its temporal validity, the requirements for recovery, and any applicable cost or budget constraints. Each judgment is linked to the documents that support it, so that the final decision and its underlying basis can be examined separately. We have tested this formulation through independent human annotation across multiple document-based scenarios, followed by a separate review of the supporting decisions and recovery and temporal requirements. The results so far show that people can apply the decision rules consistently and trace their judgments to specific evidence. We use this formulation to construct a benchmark with controlled changes to the underlying documents, cases where the available information does not support a definite decision, and checks for shortcuts based on candidate position or surface wording. The benchmark examines whether language models can distinguish actions that can be carried out from actions for which the available evidence and authority provide sufficient grounds to proceed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.