PureMem: Learner-Relative Memory Curation for Reliable Retrieval-Augmented Reasoning
Abstract
Retrieval-augmented language models commonly assume that relevant and correct context is useful. This assumption can fail when retrieved memory is redundant, difficult for the target model to exploit, or induces an incompatible reasoning strategy. We introduce PureMem, a learner-relative memory framework that asks not only what is relevant, but what a frozen language model can reliably use. PureMem constructs a compact reasoning memory bank from examples in the model's empirically estimated zone of proximal development (ZPD), and retrieves memory through three stages: candidate recall, structured procedure matching, and LLM-based applicability verification. A conservative injection policy falls back to the base model when no candidate is judged applicable. On BBEH, PureMem improves accuracy from 72.75% to 74.78% while retaining only 26.4% of the full memory bank. Counterfactual placebo interventions show that the verifier separates relatively beneficial memory from harmful memory: injected memory improves accepted cases but significantly degrades rejected cases. On RealMemBench, the three-stage retrieval harness improves NDCG@10 by 0.097 over the strongest executed embedding baseline and also improves downstream answer quality. These results distinguish a model-specific memory-curation procedure from a reusable retrieval architecture, while showing that injection policies must be adapted to the target memory domain. Code and data are provided as anonymised supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.