acceptodds
Under review as a conference paper at ICLR 2027

MIXTURE OF KV GROUPS: TRADING KV-CACHE CAPACITY FOR DECODE BANDWIDTH

Abstract

KV cache memory bandwidth has emerged as a major bottleneck to LLM inference and RL training, especially at long contexts. Grouped-query attention (GQA) has become a widely used method to accelerate decoding by reducing KV-cache traffic, but aggressive KV sharing can impair long-context quality. We study Mixture of KV Groups, which expands KV capacity while routing each token to a small subset of KV groups. Every group retains the complete sequence history, trading additional cache storage for sparse KV memory access. We train 220M-class and 3B hybrid language models from scratch, jointly learning KV representations and routing. At 220M scale and 3B parameter scales we find that Mixture of KV Groups reliably improves on the long context performance of GQA models while preserving much of their decode efficiency. We also develop an attention masking scheme to enable parallel sparse training with FlexAttention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.