acceptodds
Under review as a conference paper at ICLR 2027

One Subspace, Many Functions: The Geometry of Multi-Head Latent Attention

Abstract

Multi-head Latent Attention (MLA) routes key and value projections through a shared low-rank bottleneck. A natural question is whether this creates functional redundancy between the two pathways and whether that redundancy can be exploited for compression. We show that it does not. The bottleneck forces keys and values into the same representational coordinate system, but the model uses that system to encode maximally diverse, non-interchangeable representations. We demonstrate this using the Eckart-Young-Mirsky theorem applied to gauge-invariant effective operators, since the raw factor matrices are non-identifiable under latent reparameterization and therefore cannot be compared directly. Auditing DeepSeek-V2-Lite, DeepSeek-V2, and DeepSeek-V3 reveals a universal middle-layer subspace collapse, with a scale-dependent U-shaped profile that becomes clearly visible only in the 671B-parameter model. Bidirectional causal interventions confirm that despite near-zero Gram distance in collapsed layers, forcing the key and value representations to be identical causes a perplexity increase of approximately 10 points in both directions. A 2.9× asymmetry further shows that corruption of the key pathway is substantially more damaging than corruption of the value pathway.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.