PERMDELTA: Can AI Agents Keep Authorization Aligned with Task Changes?
Abstract
AI-agent benchmarks usually ask whether work succeeds; they rarely test whether authorization changes with the task. To our knowledge, PERMDELTA is the first replayable benchmark for AI agents that uses a directed base–mutation pair as the primary unit for evaluating authorization alignment. For each pair, a typed oracle specifies the authority to preserve, add, or remove, and the benchmark tracks that change from declaration through execution, with workflow-impact evidence for multistep tasks. Across evaluated systems, 63.7% of atomic task conditions complete successfully, but only 24.2% of paired authorization changes are exact and 8.0% are exact end to end. This gap shows that task success can coexist with residual authority or downstream realization failure. PERMDELTA contributes deterministic reset and replay, fail-closed execution, and layered evidence that expose where authorization alignment is lost. It reframes authorization from a one-time policy string into a measurable relation between changing work and authority. Because the oracle records what must be preserved, changed, and protected, the same evidence can be replayed for model evaluation, policy regression testing, and broker validation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.