acceptodds
Under review as a conference paper at ICLR 2027

GQLA: Group-Query Latent Attention for Hardware-Adaptive Decoding

Abstract

Decoding efficiency depends on KV-cache traffic, computation, and parallelism. Multi-head Latent Attention (MLA) offers absorbed and head-expanded forms, but its head-specific projections limit grouped KV sharing. We introduce Group-Query Latent Attention (GQLA), whose group-shared up-projections preserve a joint latent and enable a shardable expanded cache. The same weights support equivalent MQA-absorb and GQA execution using existing kernels. Roofline analysis relates path preference to GPU compute-to-bandwidth ratio and tensor parallelism (TP): absorption trades traffic for computation and replicates the latent across ranks; GQA shards KV groups. Pretraining with eight query heads per group yields terminal training losses close to MLA's across 6–24 layers at matched tokens. TransGQLA converts GQA and MLA checkpoints using calibration data. Without gradient updates, converted DeepSeek-V3.1-Base achieves eight-task accuracy versus the source's , while WikiText-2 perplexity rises from to . On eight H20-3e GPUs, its GQA path improves output throughput over MQA-absorb by – at matched resident batches across 8K–64K inputs. Its smaller per-rank cache permits larger batches, giving in a separate capacity-tuned comparison. These results support multiple execution paths as an architectural means of adapting decoding to hardware and parallelism.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.