MemCtrl: Do users have control over their agent's memory?
Abstract
As everyday tasks are increasingly delegated to AI agents, large amounts of personal information are disclosed and stored in memory to personalize future assistance. This raises a fundamental question: **do users actually have control over their agent's memory?** Prior work largely focuses on improving memory retention and personalization, implicitly assuming that everything should be memorized whenever capacity permits. We instead start from the premise that users may want some information forgotten while keeping other information in memory. We observe such behavior in the wild: in real-world user–ChatGPT conversations, users explicitly request that specific information be removed or disregarded, which we term *Memory Control Instructions*. Motivated by this observation, we introduce **MemCtrl**, a benchmark for evaluating memory control across three common settings: Incognito (do not store newly disclosed information), Forget (remove previously stored information), and Withhold (temporarily disregard stored information). We evaluate six state-of-the-art memory systems and two production chatbots (ChatGPT and Claude). No system reliably follows all three types of instructions: high utility consistently comes at the cost of low control, and even ChatGPT and Claude surface restricted information in 43.6%–62.4% of controlled cases on average. Through thorough analysis and targeted interventions, we trace these violations to failure modes at four levels and distill our findings into design guidelines for faithful memory control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.