acceptodds
Under review as a conference paper at ICLR 2027

PerDocBench: Evaluating Cross-Source Evidence Integration in Personalized Memory Systems

Abstract

Personalized memory agents are increasingly expected to integrate a user’s conversations with private documents, yet most evaluations still rely on homogeneous chat streams that can reward benchmark-specific retrieval shortcuts. This makes progress difficult to interpret: strong performance may reflect adaptation to benchmark artifacts rather than robust use of heterogeneous personal evidence. We introduce PerDocBench, a cross-source benchmark for evaluating whether memory systems can connect conversational cues with document-grounded facts in persona-specific scholarly contexts. Unlike chat-only benchmarks, PerDocBench targets questions that require evidence to be identified, retained, and integrated across structurally different sources. Experiments with 16 representative memory methods under a shared backbone reveal a negative-utility pattern: averaged over three runs, 15 of 16 methods recover the requested fact less often than the backbone. These results suggest that current memory systems often degrade cross-source evidence integration in scholarly document-centered personalization. An anchor-reordering ablation yields a median L1 gain of 7.1% across systems and motivates anchor-first, evidence-preserving memory construction as a future design direction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.