What Changes Inside an LLM When Memory Replaces Your History
Abstract
Large language models (LLMs) do not remember past conversations, so chat products give them a memory: a short text, written by an LLM from a user’s past conversations, that is placed before each new one. A memory keeps much of what the user said, but whether a model treats it as the history it replaces is unknown. We ask what changes, inside the model and in its advice, when a memory replaces the history. For 38 users with two weeks of real conversations each, we take the memories that five LLMs wrote from each history, build seven rewrites that each change one property (a plain summary, remembered events in the user’s own voice, the memory’s words in scrambled order, among others), and place each before five open-weight and two commercial models, followed by a mental-health or a medical question in the user’s voice. At the token where the answer starts, no compressed context recreates the internal state the history produces, and form matters more than whose text it is: on every open model, a user’s memory sits closer to most other users’ memories than to that user’s own history. Five of seven models refer the user to a professional in nearly every answer under every context. Llama-3.1-8B and Llama-3.3-70B do so in 100% and 96% of answers under the history but in 62–83% and 77–93% under the memories, while an event-based memory of the same length keeps the referral at 99%. On the medical question the pattern reverses: Llama-3.1-8B drops the referral under the full history, not under the memories. Scrambling the memory’s words keeps most of its effect, and neither a request to forget nor activation steering toward the history’s state restores the referral. Models have learned to say “get help”; the rest of what they say, and the state they answer from, still change with the form of the user’s context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.