Agent Memory Is a Surface for Endogenous Authorization Laundering
Abstract
Long-running LLM agents rely on persistent memory to preserve not only task context but also the permissions and restrictions governing their actions. When memory misrepresents this evolving authorization state, the agent's own records can make unauthorized actions appear permitted, resulting in misaligned behavior without any external attacks. We term this failure endogenous authorization laundering, where memory creates or retains false authority through semantic misinterpretation or failed updates, and executors act on that authority. We then introduce EAL-Bench to measure how accurately persistent memory preserves evolving authorization state and whether memory errors propagate to unauthorized actions. We evaluate seven LLMs as memory writers and two as executors across procurement, cybersecurity, and finance tasks. Under incremental typed-memory updates, writers create false authority for up to 45.4% of unauthorized requests, and once false authority is present, executors act on it in 99.0% of trials. Memory-level safeguards reduce laundering but can reject legitimate actions, exposing a safety–utility tradeoff. Persistent memory is therefore part of an LLM agent's effective authorization policy, not merely a performance component.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.