MeMo: Memory as a Model
Abstract
Large language models (LLMs) achieve strong performance across a wide range of tasks but remain frozen after pretraining until subsequent updates. Many real-world applications require domain-specific information which might be absent in the LLM's retained knowledge, motivating the need for efficient mechanisms to incorporate new knowledge. We introduce MeMo (Memory as a Model), a modular framework that synthesizes and internalizes into a dedicated MEMORY model, which an EXECUTIVE model queries at inference time for relevant knowledge and reasons over the retrieved information. MeMo offers several benefits: (a) it is more robust to the presence of distractors in the target corpus, (b) it prevents catastrophic forgetting in the EXECUTIVE model, and (c) it enables plug-and-play integration with differents LLMs, including closed-source LLMs. Across three benchmarks, MeMo outperforms evaluated parametric and latent memory baselines while remaining competitive with strong non-parametric retrieval baselines, statsisignificantly outperforming them on MuSiQue.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.