acceptodds
Under review as a conference paper at ICLR 2027

NibbleMLA: Dual-Axis FP4 Attention over a Single Latent Cache

Abstract

Multi-head latent attention (MLA) reduces the key-value cache to a single latent per token that serves as both key and value. What remains is numeric: how many bits store that latent, and in what precision the two attention GEMMs read it. FP4 latent caches already exist, but they dequantize to BF16 before attention. We show that this is forced by a structural mismatch rather than by accuracy: block-scaled FP4 tensor cores accept scale factors only along the contraction dimension, whereas the shared latent is contracted along channels by QK and along tokens by PV. We prove that a scale written once per token can be consumed by both GEMMs only if it is a per-token scalar times a static channel vector—precisely the layout that suits raw latents worst, because their channel magnitudes are highly uneven. NibbleMLA resolves this tension by changing the basis instead of the scale: a fixed -dimensional Hadamard co-rotation, which costs no storage, equalizes the channels so that one E4M3 scalar per token suffices. The resulting -byte-per-token cache ( FP8) feeds both attention GEMMs on FP4 tensor cores without dequantization. On DeepSeek-V4-Flash, FP4 attention matches BF16 attention computed on the same cache (output cosine in all layers); end-task accuracy is statistically indistinguishable from the FP8 cache on GSM8K, AIME, and LongBench-v2 and within points on -task RULER at K; and the decode-attention kernel is – faster than the FP8 kernel and, on average, faster than the official FP4 kernel, which dequantizes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.