Frozen transformers can reuse, compose and update state memory
Abstract
A transformer that has read a context carries its computation on that context in every layer. We show that a frozen transformer can answer from a context it no longer has when it keeps one vector per position from a single early layer, the residual activations after its second decoder layer, and reads them back by placing them at placeholder positions of a new prompt and recomputing every layer above; we call these vectors state memory. On multi-round coreference resolution (MRCR) conversations of up to 136,577 tokens, Qwen3.8-27B discarded the text and its entire key/value cache, rebuilt the other 62 of its 64 layers from its bf16 memory and scored 75.5 against 75.0 for full history (the context read as tokens); 96 of 115 greedy answers were identical (83–84% in every range of lengths). Four-bit memory scored within 0.35 F1 of full history on all 150 MultiFieldQA-en questions of LongBench with Granite 4.2 8B and Qwen3.8-27B, and on 128 synthetic handbooks bf16 memory stayed within one percentage point of full history on Granite 4.2 8B, Qwen3.5-27B and Qwen3.5-9B, against 25.5–26.2% with neither. The same frozen model also operated on these memories. On Qwen3.8-27B, four-bit memories of DROP passages and of numerical questions, each written without the other, passed 41 of 52 tests that require the right answer under all four passage–question pairings, as many as full history. Revised ten times by regenerating its text from the memory and a counterfactual correction, a bf16 memory of SQuAD paragraphs scored the same as the corrected full history on every question, and a four-bit memory revised twenty times answered 96.5% of questions against 97.9%. Finally, the model read episodes of an environment’s tool use and wrote a note; a reader holding only the note’s states executed the environment’s tool-call convention on 213 of 256 new requests on Granite 4.2 8B and 195 on Qwen3.8-27B, against 235 and 206 with the episodes in context and 0 from a note written without them. On Granite 4.2 8B, a policy memory written the same way decided 170 of 256 new records exactly, against 130 with the episodes in context, and policy and procedure memories acquired in different environments executed 438 of 1,024 composed workflows, against 202 with both logs in context.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.