acceptodds
Under review as a conference paper at ICLR 2027

Does Memory Construction Really Help the Long-Term Memory of AI Assistants?

Abstract

AI assistants need long-term memory to answer questions about users' past experiences. Existing methods reorganize these experiences into memory without considering the queries users will ask, but we find that their advantages diminish as evaluation spans more diverse benchmarks. In contrast, directly retrieving original records (Dense) remains competitive. To understand why, we keep the evidence and answers fixed and change only how questions are phrased. The resulting counterexamples show that rephrasing can turn a memory method's advantage over Dense into a disadvantage. This suggests that users' questions can play a decisive role in whether memory construction helps. Our theoretical analysis further shows that, under stated assumptions, no memory construction can guarantee a universal advantage over Dense without assumptions about the user's query distribution. This motivates Dynamic Retrieval Entrances for Agent Memory (DREAM), which preserves original records while adapting the routes used to retrieve them. As users ask questions, DREAM activates retrieval entrances that reflect their recent needs, allowing an initially retrieved record to lead to additional relevant evidence. Across four benchmarks, DREAM improves average Judge accuracy over Dense by 4.2 percentage points with Qwen3-VL-8B and 6.2 points with Gemma4-12B.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.