Memory Navigates, the Record Decides: Auditable Evidence for Million-Token Agents
Abstract
Even when a million-token history fits in the context window, an agent cannot attend to all of it at every step, so it answers from compressed memory. Memory systems that answer from model-written summaries and facts can use outdated or contradictory information and cannot tell when evidence is missing. We observe that compiled memory can afford to be lossy if it only points to source lines that the agent re-reads, so we use it only to navigate to evidence and let the re-read record decide the answer. NavMem, a training-free compiler, builds a multi-resolution continuity view of the history and a dynamics view that tracks updates, conflicts, and event order. Code stamps each item with the address of its source lines, so every item is auditable, and the answer protocol asks the agent to check each claim against the re-read lines and to abstain when they do not support it. When every system answers with Qwen3.5-122B, NavMem scores 0.715 macro on BEAM's 1M tier, 0.100 above the strongest baseline, Mem0 (paired 95% CI [0.076, 0.125]), and 0.755 accuracy on LoCoMo, 0.066 above LIGHT ([0.051, 0.083]).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.