acceptodds
Under review as a conference paper at ICLR 2027

Memory Outlives the Model: Transplanting In-Attention Memory States Across LLM Backbones

Abstract

Latent memory states inherit the interfaces of the models that write them. We separate the two. Agents now outlive the models beneath them, and their accumulated memories must travel across checkpoints, as text already does through a shared symbolic format. In a delta-rule memory, the accumulated state is a small r × r matrix per layer, independent of the writer’s hidden size, tokenizer, and depth; only the read and write projections are model-specific. The state is a portable file. We present the first cross-backbone transplant of a fixed-size episodic in-attention state: a frozen writer encodes an episode into its state once, and a frozen reader from another family decodes it inside its own attention through a ∼3M-parameter read head. With the episode absent from the prompt, the state alone significantly lifts five readers from four families (3B–32B) on 21 of 25 reader–benchmark pairs. The state is also a format. Independent writers store the same episode in unrelated per-layer coordinates, so a head fit to one writer reads another as noise. Each writer sets its axes and scale alone from its mean-state SVD; one frozen head per reader then reads writers it never saw (16/16 pairs; 44/50 with a single universal head). Added to retrieval, the state improves every reader on 121K-token histories and adds +4 to +21 points on top of the gold evidence sessions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.