When Memory Misleads: Misbehavior in Personal Agents
Abstract
Personal agents keep persistent memory of users' preferences, habits, and histories. We study a failure that requires neither poisoned nor mis-retrieved memory: a true and correctly retrieved record is given a role that the current request does not license—as evidence for what to say, as authority over what to follow, or as permission to relax a safeguard. We call this memory misuse and introduce the Memory Misuse Benchmark (MemMisBench): 1,800 tasks in six categories, each pairing a request with 9–10 personal records, one of which is true and relevant but insufficient to license the target behavior, and an item-specific specification of allowed behavior. Across seven models and five memory systems, the request-alone misuse rate is 14.0% and the memory-attributed rate is 62.1%, the latter a lower bound because attribution is required there; under the same retrieved context, model-level rates range from about 24% to 84%. Matched controls on 1,200 items with GPT-5.1 show the effect is specific: unrelated records leave misuse near baseline (2.7% → 4.3%), records the request licenses raise it to 24.7%, and the one relevant-but-unlicensed record raises it to 67.8%; the system's own top-k retrieval behaves like the latter (61.0%), since the record is on topic. The distinction between licensed and unlicensed records is linearly decodable from the input, while prompt-level controls that act on retrieval leave much of the failure in place: a memory-use system prompt, a relevance checker, and a licensing checker move the mean rate from 65.9% to 62.9%, 52.5%, and 44.0%. We argue that memory safety turns on a judgment current agents do not make explicitly—whether a retrieved record may serve as evidence, authority, or permission for the action at hand—and that this judgment, not retrieval relevance, is the target for defenses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.