MemAny: Plug-and-Play Generative Latent Memory for Any Model
Abstract
Knowledge acquired through model training is typically tied to the model that learned it, requiring repeated training as new models emerge. We introduce MemAny, a framework that turns adapted knowledge into plug-and-play generative latent memory that can be trained once on a host model and reused across different recipient models. MemAny routes adaptation into an external memory weaver while keeping the host model frozen, and expresses its generated latent tokens as linear combinations of each recipient’s own token embeddings. This latent superposition enables knowledge transfer across tokenizers through coefficients aligned by shared vocabulary anchors, without updating either the recipient or the learned memory. Experiments combining MemAny with various training algorithms, e.g., supervised fine-tuning (SFT), reinforcement learning (RL), and on-policy distillation (OPD), demonstrate that memories learned with small host models yield transferable gains in mathematical/logical reasoning, code generation, and tool-use. In particular, OPD-trained memory from Qwen3-4B-Base transfers across model families and tokenizers to MiMo-V2-Flash-Base (309B) and HY3-Preview-Base (295B), improving their AIME24 accuracy by up to points. These results establish portable generative memory as a practical interface for sharing acquired capabilities, paving the way for their efficient and rapid dissemination across an evolving model ecosystem.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.