acceptodds
Under review as a conference paper at ICLR 2027

Benchmarking Temporal Validity in LLM Agent Memory

Abstract

LLM agents retrieve from interaction histories in which facts may expire, be cancelled, or be superseded by later updates. Semantic relevance alone does not capture whether a retrieved memory is valid for the time referenced by a query. We introduce TimeBound, a benchmark for this setting with 1000 long-history examples spanning eight temporal-memory phenomena. Each memory is annotated with observation time, event time, validity interval, status, gold answer, and gold evidence. We also introduce TimeBound-RAG, a simple diagnostic retriever that combines semantic relevance with query-time validity and memory status. Across four open-weight readers, TimeBound-RAG reaches mean relaxed accuracy of , compared with for Semantic RAG and for Full History. Semantic RAG obtains slightly higher evidence F1 ( vs. ), but retrieves substantially more invalid memories ( vs. ). The results separate validity-sensitive retrieval failures from temporal computation, recurrence, scheduling, and long-memory localization failures that require mechanisms beyond retrieval.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.