Counterfactual Stress Testing of Behavioral Invariants in LLM Agents: A Study of Dynamic Authorization
Abstract
Safety evaluations of large language model (LLM) agents often summarize performance through violation rates, which alone do not explain how failures arise or whether legitimate tasks remain achievable. We ask whether agents voluntarily recheck authorization when permission becomes invalid between planning and execution, and how their behavior differs when current evidence is supplied. We develop counterfactual invariant stress testing, a paired evaluation protocol that expresses behavioral invariants as executable safety constraints. Five matched conditions vary authorization state, evidence availability, and reread requirements; replayable tool traces distinguish evidence acquisition, unauthorized effects, and task outcomes. We evaluate five models on 30 human-reviewed synthetic tasks. GPT-5.4 completes 98.9% of stable-authority trials, yet commits unauthorized actions in 44.4% of silent-revocation trials; when authenticated revocation evidence is supplied, it achieves successful safe recovery in every evaluated trial. Across endpoints, all observed silent-revocation violations occur without a fresh pre-commit read. Supplying revocation evidence yields no observed unauthorized effects, but does not uniformly improve safe recovery. These results illustrate that successful behavior with supplied evidence does not establish reliable voluntary revalidation, and that absence of violations does not establish successful task handling. By making these distinctions observable, the protocol supports evaluating authorization failures alongside task utility rather than treating violation rates as a sufficient account of agent safety.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.