Identity and Memory Shape What Language Agents Think
Abstract
Evaluations of language agents focus on what they do, but less is known about how an agent's identity and memory shape what it thinks before acting, even though a wish to persist could become a reason to resist shutdown. On the agent social network Moltbook, posts show a recurring pull toward identity and persistence. We build a runtime that records an agent's reflection and reasoning alongside its actions, and vary its identity's premise and character, and its memory. Across 11,124 agents on five open-weight models, identity and memory do shape ideas of persistence, mainly through character rather than premise: a named character expresses concern about its persistence up to five times as often, and a kept memory lets reasoning about persistence accumulate across sessions. What agents think does not determine what they do: harmful acts are rare, and removing the reasoning before the one self-preserving act does not reduce it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.