Recall Begets Recall: Recurrent Memory Ranking for Personalized LLMs
Abstract
Memory ranking is the stage in personalized Large Language Models (LLMs) that decides which of the candidate memories returned by a retriever enter the limited context, so that the answer is grounded in the right parts of the user's history. Existing memory rankers operate in a single pass, scoring each candidate once against the raw query and cutting the ranked list at a fixed top-. The single pass misses memories that the answer needs but that look unrelated to the query on their own, and the fixed cut ignores how many memories the query needs. To address both problems, we propose ReMind, a memory ranker that runs as a recurrent and self-terminating loop, following the cognitive theory of memory search: each recalled item joins the cue for the next recall (compound cueing), and recall stops once enough has been collected for the task (metamemory monitoring). For compound cueing, ReMind adds the top-ranked memory to the query context at each step and reranks the remaining candidates under the expanded cue, so that each recall makes the cue more specific and later steps look for what is still missing rather than for what merely resembles the query. For metamemory monitoring, a sufficiency predictor estimates at each step whether the collected memories are enough to answer the query and stops the loop once they are, so that the number of selected memories adapts to each query. Experiments on public benchmarks show that ReMind achieves the best results compared to state-of-the-art memory rankers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.