When Relevant Memories Are Wrong: A Controlled Reliability Audit of Episodic Memory Retrieval
Abstract
Episodic memory systems retrieve prior episodes by semantic similarity, but relevance does not guarantee a correct outcome. We construct paired GSM8K memories that preserve the task and most of the worked derivation while changing the final conclusion, and evaluate retrieval and downstream use under matched memory contexts. Across all 1,319 test questions, Qwen3-Embedding-4B ranks the verified episode first in 40.3% of final-only pairs and 66.4% of structured pairs; BGE-M3 reaches 32.1% and 39.8%, respectively. For BGE-M3, paraphrasing changes selection by at most 0.2 percentage points, whereas removing the duplicated question changes Good@1 by 9.3–12.6 percentage points. In a frozen response bank, verified memory raises accuracy from 93.0% to 99.0%, while corrupted memory reduces it to 56.5%. Semantic Top-1 reaches 81.0% versus 77.5% for matched random selection; the paired 95% CI includes zero. Natural incorrect-solution and temporal-evidence evaluations show that retrieval relevance, episode correctness, and generator utilization should be reported as distinct quantities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.