BetweenUsBench: Resonance Requires Dyad-Specific Memory, Not Mere Recall
Abstract
Long-term conversational memory supports both task execution and companionship. Beyond retaining facts, companionship requires understanding a user’s private meanings and context-dependent response expectations, supporting a sense of resonance: this agent understands me. Retaining conversational evidence does not ensure that an agent can infer and apply these mappings. We introduce BetweenUsBench, a benchmark for this capability, which we call dyad-specific memory. Its two dimensions are idiosyncratic reference, interpreting private codenames, in-jokes, and metaphors, and interactional rapport, following response expectations learned from prior interactions. Each memory system ingests a complete multi-session history and answers a subsequent query. An LLM judge assesses the reply against the history-supported mapping. History controls, a rapport-specific profile baseline, and supplementary diagnostics examine specificity, preservation, retrieval, and utilization. Across three memory backbones with a fixed reader, the best memory-method pass rates range from 20.42% to 25.42% for private reference and from 24.17% to 28.75% for rapport. Failures persist even when sufficient evidence reaches the reader. BetweenUsBench makes these limitations measurable, connecting response quality to how remembered experience is recovered and used.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.