acceptodds
Under review as a conference paper at ICLR 2027

ShareSVD: Sharing Key Subspaces for Low-Rank Generative Transformers

Abstract

Singular value decomposition provides a practical low-rank factorization of attention projections, reducing parameter count and computation while supporting compact latent KV caches. Existing methods typically factorize Query, Key, and Value projections independently, ignoring both their unequal sensitivity to rank reduction and cross-head subspace redundancy. Our systematic study across vision-language models (VLMs) reveals a consistent asymmetry: Value is the most fragile projection, whereas Key is the most robust. Motivated by this observation, we introduce ShareSVD, which adaptively groups geometrically compatible Key heads and shares a low-dimensional input basis within each group while retaining head-specific output factors. This sharing reduces the parameters required for Key factorization, enabling more low-rank directions to be retained in the more sensitive Query and Value projections. At a retained QKV ratio of 0.3, ShareSVD improves average accuracy over the competitive baseline by 8.57% across three VLMs and two benchmarks, while reducing low-rank KV-cache memory by 14.7%. Consistent perplexity improvements on dense and MoE Qwen3 models further demonstrate that ShareSVD generalizes beyond multimodal architectures. Code is available at https://anonymous.4open.science/r/ShareSVD-F0E0.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.