acceptodds
Under review as a conference paper at ICLR 2027

When Personal Memory Has No Single Answer: Evaluating LLM Agents Facing Underdetermined Conflicts

Abstract

Large Language Model (LLM)-based agents increasingly retain personal memory across sessions, but such memory can contain inconsistencies involving context-dependent preferences, evolving behaviors, and conflicting sources. When a query omits the context, time, or source authority needed to interpret these inconsistencies, treating one memory as definitive produces unjustified, overconfident actions. Existing benchmarks reward designated outcomes, overlooking whether agents recognize underdetermination, preserve alternatives, seek missing information, and express evidence-compatible behavior. We introduce Testing Agents' Navigation of Genuine, Latent, and Entangled Memory Conflicts (TANGLE), a controlled benchmark for evaluating cognitive behavior when agents face underdetermined conflicts in personal memory. TANGLE comprises 541 fully synthetic instances that span 40 diverse personas and three conflict types: Context-Partitioned Conflict (CPC), Behavior-Oscillation Conflict (BOC), and Source-Contradiction Conflict (SCC). We evaluate matched oracle and pipeline tracks, using curated memory and system-retrieved memories from multi-session dialogues, respectively. Results show that explicit conflict recognition is more frequent than calibrated-action or targeted-clarification behavior under the rubric; retrieved memories can also omit relations needed for downstream reasoning. Fixed resolution rules show structural limitations; we therefore introduce Conflict-Aware Action Policy (CAAP) as a two-stage prompting-based policy baseline that selects evidence-grounded actions case by case. Together, TANGLE and our findings frame conflict handling as recognizing underdetermination, retaining conflicting evidence, and selecting an evidence-compatible response.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.