acceptodds
Under review as a conference paper at ICLR 2027

Towards Long-Term Memory with Small Language Models

Abstract

As personal AI assistants proliferate, they are expected to become capable of complex, long-horizon interactions while preserving the privacy of increasingly personal information. A central challenge in achieving this form of personalized intelligence is long-term memory. An assistant cannot genuinely understand a person from isolated sessions, it needs a continuous thread of experience that accumulates over time, where someone has been, whom they have spoken to, what they have been working on, and the context connecting these experiences. This requirement for persistent memory is tightly coupled to privacy. Rather than transmitting an ever growing record of personal experience to a remote server, data center, or other external system, memory can remain private by being stored and processed locally on the device where the assistant operates. Such a design, however, places a fundamental constraint on the underlying model, mainly the assistant's memory must be small enough to run on a phone, a pair of glasses, or a robot while retaining the capacity to maintain and recall over a persistent stream of experience. Yet every long-term memory result we are aware of is reported at model scales an order of magnitude beyond that budget, and whether a small model can carry such an ability at all is, to the best of our knowledge, underexplored. In this paper we propose an approach to long-term memory built purely on small language models, in which memory lives in intrinsic recurrent state rather than in a growing attention window. We show that small models combined with linear attention and the right training strategy offer strong recall, and that because a bounded state has no key-value cache to grow, such a model can ingest conversations far longer than a large transformer can admit at all. We evaluate against open-source counterparts such as Gemma-3-27B-it, Qwen2.5-32B-Instruct and Qwen3-30B-A3B, on the BEAM long-term memory benchmark and find our approach competitive with models an order of magnitude larger.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.