AGENT MEMORY SCORERS ARE UNRELIABLE ON DENIAL OUTPUTS
Abstract
Agent memory is important for enabling agents to retain, retrieve, and effectively use information from long-term interactions. An LLM agent's memory capability is typically evaluated via *questions*: a fixed *reader* answers each question based on the *memory context* returned by the system under test, and a *scorer* evaluates the answer alone. This work claims that this scoring mechanism is severely biased on questions about an item the recording never contained, whose credited answer is a denial, *no such thing exists*: a claim about the whole recording that no scorer checks. Such questions make up about a quarter of LongMemEval-V2 and LoCoMo. On LongMemEval-V2, adding a simple note, *the recording is complete, so what is not listed does not exist*, to the memory context of the released retrieval baseline raises its denial accuracy from 16% to 72% without changing the retrieval result. The same note makes it deny 55 objects that do exist, up from 1, and the score never notices: a false denial and an honest abstention both score **0**. We then propose three specific mitigation strategies, one for each main component of the agent memory evaluation pipeline: *questions*, *reader* and *scorer*. Extensive experiments show that once the scorer checks each denial against the recording, at most 4 of the baseline's 92 credited denials stand. Read off the answer, a denial score measures who asserts completeness; read against the recording, who can show it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.