acceptodds
Under review as a conference paper at ICLR 2027

RAQ-MLA: Role-Aware Quantization for Multi-Head Latent Attention Caches

Abstract

Multi-Head Latent Attention (MLA) has been adopted by recent frontier large language models. Instead of per-head keys and values, MLA caches a compressed KV latent that is up-projected into keys and values at decode time, together with a narrow positional key that carries position information into the attention score. Existing KV-cache quantization methods are designed around the keys and values of standard attention and do not account for what MLA actually caches, which costs accuracy at low bit widths. We analyze the MLA cache both in activation distribution and in how quantization error propagates to the attention output, and show that the two cached tensors play different roles: the latent mainly supplies the values, whereas the positional key dominates the attention score. We therefore propose RAQ-MLA (Role-Aware Quantization for Multi-Head Latent Attention Caches), which fits an invertible transform to each cached tensor against the error that tensor delivers to the attention output: attention-weighted value error for the latent and score perturbation for the positional key. Without training, RAQ-MLA quantizes the full MLA cache to about bits per element, a increase in theoretical KV-cache capacity, while remaining nearly lossless on reasoning and coding tasks across four MLA models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.