ARAS: ADAPTIVE RANK-ALLOCATED SHARED SUBSPACES FOR LLM COMPRESSION
Abstract
Deploying foundation models across diverse hardware platforms requires reducing model size without substantially degrading model quality. Existing low-rank compression methods often rely on global approximations, complete rows or columns, predefined spatial blocks, or fixed sharing patterns. We observe that Transformer weight matrices contain finer low-dimensional structure: spatially distant fine-grained weight blocks can often be represented by the same subspace. Based on this observation, we introduce Adaptive Rank-Allocated Shared Subspaces (), a post-training compression framework that learns which fine-grained weight blocks should share a low-rank basis, independent of their locations within the matrix. further allocates rank adaptively using a small calibration set under a global storage budget. Across OPT models from 125M to 30B parameters and Llama models from 1B to 8B parameters and compression levels, reduces the overall model size by 67% while maintaining competitive perplexity and consistently outperforms the evaluated low-rank baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.