KVPath: Connecting Key-Value Representations Across Transformer Layers
Abstract
Transformers typically cache a separate key-value (KV) representation for each token at every layer. We hypothesize that most of a token's KV representation can be reused across neighboring layers, while the change from one layer's KV state to the next can be represented by a low-dimensional delta. Motivated by this hypothesis, we introduce KVPath, a new KV representation paradigm that connects KV states across network depth during pre-training. KVPath organizes these states as a path, with one component remaining unchanged across layers and a low-dimensional component varying by layer. KVPath supports major attention architectures, including grouped-query attention (GQA) and multi-head latent attention (MLA). Experimental results across multiple model scales and context lengths show that KVPath achieves better performance than GQA and cross-layer attention (CLA) at matched model sizes and KV-cache budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.