acceptodds
Under review as a conference paper at ICLR 2027

Counterfactual Stress Testing of Behavioral Invariants in LLM Agents: A Study of Dynamic Authorization

Abstract

Safety evaluations of large language model (LLM) agents often summarize performance through violation rates, which alone do not explain how failures arise or whether legitimate tasks remain achievable. We ask whether agents voluntarily recheck authorization when permission becomes invalid between planning and execution, and how their behavior differs when current evidence is supplied. We develop counterfactual invariant stress testing, a paired evaluation protocol that expresses behavioral invariants as executable safety constraints. Five matched conditions vary authorization state, evidence availability, and reread requirements; replayable tool traces distinguish evidence acquisition, unauthorized effects, and task outcomes. We evaluate five models on 30 human-reviewed synthetic tasks. GPT-5.4 completes 98.9% of stable-authority trials, yet commits unauthorized actions in 44.4% of silent-revocation trials; when authenticated revocation evidence is supplied, it achieves successful safe recovery in every evaluated trial. Across endpoints, all observed silent-revocation violations occur without a fresh pre-commit read. Supplying revocation evidence yields no observed unauthorized effects, but does not uniformly improve safe recovery. These results illustrate that successful behavior with supplied evidence does not establish reliable voluntary revalidation, and that absence of violations does not establish successful task handling. By making these distinctions observable, the protocol supports evaluating authorization failures alongside task utility rather than treating violation rates as a sufficient account of agent safety.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.