Last Layer Key-Value Cache is All You Need
Abstract
The key-value (KV) cache of an auto-regressive transformer grows with both sequence length and network depth: every layer stores its own keys and values. Cross-layer KV cache sharing lets several layers read one layer's cache, but existing designs share only among a few layers, each reading a cache from a lower layer. We want the opposite: every layer reading the last layer's KV. Done at every token, this would break exact parallel training, so shares at chunk boundaries instead, as a chunk recurrence. Each chunk runs as a standard per-layer-KV Transformer and, once complete, writes only its last-layer KV cache, which every layer of every later chunk reads through ordinary attention. Within a chunk, training keeps a standard Transformer's parallelism. Only the most recent chunk keeps every layer's KV; every older chunk keeps only the last layer's. The remaining question is: can auto-regressive Transformers actually work from one shared last-layer cache? Across language modeling, novel view synthesis, and autoregressive video generation, experiments show that task quality holds or improves: lowers language-model validation loss and stays competitive on RULER retrieval, raises novel-view PSNR for 3D reconstruction, and maintains VBench scores for video generation. Continued pretraining converts standard Transformer weights into , which applies wherever an autoregressive model must carry long-context memory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.