Is Complex Agentic Memory Always Needed? Naive RAG Is Stronger Than We Thought
Abstract
LLM memory is widely used in long-term conversation understanding. Existing agentic memory systems often manage memories through mutation, structural organization, learned retrieval strategies, or other complex operations. Is this complexity always necessary? We study Raw-RAG, a naive RAG baseline using dense retrieval and reranking over raw history to supply evidence to a reader LLM. Across LoCoMo, LongMemEval, MEME, and MemoryAgentBench, Raw-RAG ranks first among the compared memory systems in six of seven benchmark–model settings, outperforming the strongest compared agentic memory system by 0.13–5.53 points in those settings. We then investigate why raw-history retrieval remains competitive in these settings. Evidence analysis shows that its higher coverage of necessary evidence can offset lower utilization. Useful evidence includes both answer values and the relations and dependency paths needed to derive them. Even complete access to transformed memory stores leaves an accuracy gap relative to raw evidence, highlighting the importance of preserving raw information. These findings challenge the need for complex agentic memory systems in the evaluated conversation tasks and establish naive RAG as an essential baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.