MEMORY AS A CACHE: EXACT CONTEXT REUSE AND DELETION BY CONSTRUCTION
Abstract
The KV cache of a transformer entangles every token's representation with its entire prefix: a passage encoded once cannot be reused under a different prefix, and cannot be removed without recomputing everything after it. Exact cache reuse in serving is therefore limited to shared prefixes, and systems that reuse more must approximate. We present SMem (static memory), an architecture whose context representation is a cache by construction. A block-local encoder maps each block of tokens to memory rows independently of all other blocks; a reader conditions generation on the union of those rows through cross-attention. Three properties then hold for every parameter setting: memory composes exactly at fixed block indices, deleting a block is exact and is an memory update for -token blocks, and the memory state is independent of the edit path. What the cache buys is measured directly. At the training context, under the shared training recipe, SMem still retrieves planted needles beyond any window of the trained length (exact-match accuracy – at distances 31 and 63 blocks) where a learned-position transformer, a RoPE transformer, a Block-Attention-style two-stream arm, and windowed or saturated readings of the transformers all score at most ; with two identical keys it returns the nearer one's value. Because reader computation is block-local, a fully cached context is served by computing one block alone (exact up to floating-point rounding) at a near-constant – ms against the same model's cold prefill, which grows with context; batched decode holds – fewer KV rows and runs – faster than the transformer where decode is bandwidth-bound, and end-to-end deletion beats suffix recomputation by at 512 blocks, rising to at 4096 blocks; both are – the trained length and probe the cost model rather than a served regime. What it costs is a perplexity gap of to (negative favors SMem) against a parameter-matched transformer with the same positional scheme, at 160M–1.5B on FineWeb-Edu across two training recipes and a learning-rate search; the band's lower endpoint is set by the recipe and its upper by the learning-rate search, and the recipe's change moves the transformer twice as far as it moves SMem at 160M. SMem also composes with RoPE: the composite leads both the matched transformer and SMem on every seed pair, closing – of SMem's distance to a RoPE transformer, which leads SMem by –. Context representation entangled with the full prefix is a design choice; an architecture that drops that entanglement reaches comparable perplexity and gains a cache that composes, deletes, and edits exactly and, under the shared recipe, retrieves at range where every transformer variant we test does not.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.