Investigating Tool-Memory Conflicts in Large Language Models
Abstract
Tool use is a defining capability of LLM-based agents, enabling them to act beyond their parametric knowledge and interact with external environment. Yet this reliance on external tools introduces a new vulnerability: the knowledge returned by tools may directly contradict an agent's internal memory. In this paper, we identify and systematically study this phenomenon, termed Tool-Memory Conflict (TMC), a previously unexplored class of knowledge conflict for LLMs. We find that even state-of-the-art LLMs struggle under TMC, with notable performance degradation on STEM-related tasks. We further reveal that LLMs prioritize tool knowledge and parametric knowledge differently depending on factors such as task domain and conflict severity. We evaluate existing conflict-resolution techniques, including prompting-based and retrieval-augmented methods, and show that none effectively resolve tool-memory conflicts, highlighting TMC as an open and critical challenge for building reliable LLM-based agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.