MIXTURE OF KV GROUPS: TRADING KV-CACHE CAPACITY FOR DECODE BANDWIDTH
Abstract
KV cache memory bandwidth has emerged as a major bottleneck to LLM inference and RL training, especially at long contexts. Grouped-query attention (GQA) has become a widely used method to accelerate decoding by reducing KV-cache traffic, but aggressive KV sharing can impair long-context quality. We study Mixture of KV Groups, which expands KV capacity while routing each token to a small subset of KV groups. Every group retains the complete sequence history, trading additional cache storage for sparse KV memory access. We train 220M-class and 3B hybrid language models from scratch, jointly learning KV representations and routing. At 220M scale and 3B parameter scales we find that Mixture of KV Groups reliably improves on the long context performance of GQA models while preserving much of their decode efficiency. We also develop an attention masking scheme to enable parallel sparse training with FlexAttention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.