GroundedMem: Source-Resolved Retrieval and Conservative Refusal for Long-Term Conversational Memory
Abstract
Long-term conversational memory requires recovering information from growing interaction histories. However, retrieval relevance does not imply evidential sufficiency: related memories may refer to the wrong participant, time, or event state, while missing support in a bounded retrieval set does not establish global unanswerability. Most long-term-memory systems do not explicitly report refusal quality, while those that do remain substantially weaker on this axis. We present GroundedMem, an auditable pipeline linking retrieval, source recovery, and selective answering. Compact Atomic memories serve as semantic indices; Atomic and direct-turn retrieval are fused and reranked; and selected candidates are resolved to complete, source-linked conversation turns. A controller refuses only when constrained mismatch checks fire without detected support; unresolved cases remain on the ordinary answer path. We develop the verifier prompt and hard decision policy on LongMemEval-S and transfer them to LoCoMo after normalizing candidate certificates, without LoCoMo-label tuning of the verifier. Relative to the published comparator rows included in our comparison, GroundedMem improves LoCoMo LLM-judge accuracy by 5.4 points and LongMemEval-S accuracy by 3.4 points. LoCoMo refusal F1 exceeds the reported REMem value by 9.3 points. In benchmark-internal paired comparisons, with refusal disabled and the Answer model, Judge protocol, prompt, and evidence budget fixed between arms, the complete evidence path exceeds Direct+EvoEmbedding by 2.97 points on LoCoMo and 1.60 points on LongMemEval-S. A separate paired LoCoMo ablation yields a net accuracy gain of 8.06 points from refusal, with newly incorrect answers accounting for 0.45% of answerable questions. Complete annotated source coverage nevertheless leaves substantial answer errors, motivating separate evaluation of discovery, evidence use, and answer decisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.