acceptodds
Under review as a conference paper at ICLR 2027

Can Agents Develop Private Shorthand with You? Benchmarking Dyadic Convention Formation in Long-Term Conversations

Abstract

Long-term conversational agents must do more than recall facts: they should acquire the partner-specific expressions that emerge through repeated interaction. Existing long-term memory benchmarks test information stated in dialogue rather than conventions that emerge across sessions, leaving open whether an agent can participate in their formation and later use them. We introduce PactBench, a benchmark of 200 dyadic conventions spanning inversion, anchoring, adoption and erosion. Each convention has user-compression and agent-compression histories that test whether the agent can compress a reference, accept its partner's shortened form and later produce the established form. We evaluate each required capability independently by withholding the agent turn at that session while preserving the reference formation trajectory. To prevent plausible but ungrounded replies from passing, conjunctive scoring combines a direct meaning question with capability-specific checks wherever a form has been established. Experiments with four backbone LLMs and six existing memory systems identify two failures: (i) LLM-based memory systems lose or fail to retrieve the form–meaning link, and (ii) even when given the complete preceding reference formation history, agents do not reliably compress or produce the form. We present Pact-Mem, which preserves convention-level form–meaning bindings across sessions. Across four backbone LLMs, Pact-Mem consistently improves convention production over existing memory systems by up to 31 percentage points. Pact-Mem provides a convention-aware memory module for preserving and using dyadic conventions in long-term conversations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.