Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
Abstract
Long-term memory systems typically retrieve records similar to a new query. This assumes that a needed memory will resemble the query that needs it, but world knowledge can make dissimilar texts consequentially related: a tree-nut allergy should change an answer about macarons through their almond-flour ingredient. We call this failure mode the implicit-association blind spot and introduce InMind, a 125-task, expert-verified benchmark with paired controls that distinguish storage failure, missing bridging knowledge, and retrieval failure. With the decisive memory in context, GPT-5-mini gives the required reminder on 85.6% of indirect queries. Six vector, graph, and agentic memory systems reach at most 8.0% when success requires both recalling the fact and giving the reminder, despite reaching up to 100% direct-query accuracy. A larger embedding improves recall but leaves the gap intact. Conversely, a diagnostic probe that makes memory visible before the query arrives reaches 64.8%, locating the failure in query-conditioned retrieval and motivating routing: deciding which facts must remain visible.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.