acceptodds
Under review as a conference paper at ICLR 2027

KV-Fold: Beyond Chunked Prefill with Manipulable KV State

Abstract

Long-context inference in pretrained transformers relies on a growing key–value (KV) cache that is usually treated as an internal artifact of attention. We instead view the KV cache as an explicit computational state that can be carried, modified, and transferred between computations. Standard chunked prefill is the identity case of this view: carrying the KV state unchanged across chunks is exactly equivalent, in exact arithmetic, to a monolithic causal forward pass. We prove this equivalence and verify it across several pretrained transformer families, with next-token loss matching full attention to within nats. This exact identity provides a controlled reference point for studying what happens when the carried state is changed. Without modifying model parameters, we quantize, selectively retain, re-address, and transfer KV state. The unmodified state preserves exact factual retrieval across contexts up to K tokens and chunk transitions while reducing peak working memory enough to process a context for which the corresponding monolithic forward does not fit on the same hardware. Under lossy transformations, per-step quantization reveals a graded memory–fidelity trade-off, whereas position removal is substantially less forgiving. Most strikingly, KV state can serve as a sparse message between otherwise separate inference processes. At K tokens, a consumer that never reads the original document recovers a designated fact in all trials from a shared subset of only of cache positions; matched-bandwidth random and recency controls recover none. A per-head question-aware selector succeeds even at retained positions, and RoPE-correct re-addressing allows selected state to be moved into a new positional coordinate frame without loss of retrieval. These results suggest that the KV cache is not merely stored context, but a manipulable state interface through which pretrained transformers can preserve, compress, address, and communicate information without architectural changes or fine-tuning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.