Exploiting Weight Space Symmetries for Vision Attention Compression
Abstract
Transformers have achieved remarkable success across diverse domains, yet their large model sizes incur considerable computation and latency. Accordingly, weight decomposition has emerged as an effective compression technique. However, previous methods do not consider the unique mechanism of Multi-Head Attention (MHA), where query-key (-) and value–output (-) computations are linear operations. To effectively leverage this mechanism, we propose Unilateral Decomposition for vision model compression (UniDec), a novel framework that applies decomposition to only one side of the - or - weight pairs. This approach effectively addresses approximation error at the unilateral weight level while exploiting the linear characteristics of the attention. Additionally, since --- weights exhibit varying sensitivities to low-rank approximation across heads, UniDec adaptively selects which side to be decomposed according to the rank sensitivity, thereby preserving the important information of weights. Extensive experiments demonstrate that UniDec generalizes to various transformer-based structures. Demonstrating exceptional performance even under aggressive compression (i.e., 60% MHA reduction), UniDec improves Top-1 accuracy by +34.0% / +46.3% on DeiT-Small / DeiT-S Distill while maintaining a +0.7% gain on DeiT-Base.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.