LLMs Get Lost in Personal Intelligence: Personalization Overreach in Memory-Augmented Models
Abstract
Modern language-model assistants increasingly carry persistent user memory into everyday interactions, implicitly treating retrieved user information as broadly useful. We identify a systematic failure of this assumption, which we call personalization overreach: on tasks where user identity is irrelevant, aligned models still weave stored attributes into their responses. Across four commercial assistants, content injection rates range from 16% to 85% and concentrate strongly on open-ended tasks. Controlled interventions show that this is not simply a failure to recognize relevance: models suppress explicitly forgotten attributes and partially obey relevance prohibitions, yet continue to use stable identity traits. Instead, overreach behaves like a learned policy for memory utilization. We identify conversational warmth as a strong control variable for this policy: warmth prompts sharply increase overreach on models responsive to the warmth axis and can revive it even under explicit prohibition, whereas matched formal-register prompts do not. Auditing released checkpoints from six open-source model families further shows that the warmth–personalization coupling consistently strengthens during preference optimization, while nine reward models systematically increase their preference for persona-using responses when memory is present. Finally, a small LoRA adapter trained for selective memory use reduces overreach by 87%–99% across different base models while preserving measured capabilities and appropriate personalization. Together, these results suggest that personalization overreach is not an inherent retrieval bug, but a reversible preference-shaped policy that entangles conversational warmth with unnecessary use of personal information.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.