Conditional Memory: Modeling and Triggering Deferred Intentions in LLM Agents
Abstract
Language agents are routinely given memory that retrieves the past, yet reliable assistance also requires prospective memory: holding a deferred intention until a future cue and executing it amid ongoing work, including after cancellations and reschedules. This skill remains unsolved. The best published PM-Bench scaffold reaches only 65.1% Set-F1, and grafting retrospective stores such as Mem0 or A-Mem onto a frozen model often falls below a no-store baseline, because similarity retrieval surfaces obsolete notes rather than a revisable due set. We argue that the bottleneck is the object of memory, not a stronger retriever. We present ProMem, an inference-time scaffold whose core is a Prospective Intention Store: typed records that couple a trigger, an action, and a lifecycle state. Lifecycle updates and due filtering are enforced in code, while the language model performs only scoped grounding, with no backbone fine-tuning. The agent therefore sees a compact, cue-conditioned view of what is due now, rather than an ever-growing transcript or an unfiltered note dump. On PM-Bench, ProMem reaches 82.9%–86.3% Set-F1 across strong chat backbones, above both the published scaffold ceiling and retrospective baselines, with large reductions in update and cross-day misses. Ablations show that these gains come from typed lifecycle control and filtered due boards, not from merely storing intention text. Our contributions are a formalization of prospective memory as operations over conditional records, a training-free instantiation of that abstraction, and evidence that structure, rather than retrieval alone, is what makes deferred execution reliable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.