acceptodds
Under review as a conference paper at ICLR 2027

Learning to Sketch for Low-rank Model Compression

Abstract

Transformer blocks are the cornerstones of modern deep neural networks. While energy consumption used to be dominated by training of the models, widespread adoption and intensive usage shifted it towards model inference - thus highlighting the importance of post-training model compression for transformers. Sketching describes the field of probabilistic compression that use random projections, sampled from a probability distribution. While being widely used for data compression, sketching for model compression remains widely unexplored. One key component of the sketching methods is the choice of probability distribution over sketching matrices. We propose data-aware sketching by learning the sketching distribution through training a variational autoencoder (VAE). To preserve the utility of the low-ranked model on the downstream task, we train the VAE with a carefully tailored loss function. We show that our method outperforms data-oblivious sketching and low-rank compression baselines. *For compression ratios of up to %, our compressed model outperforms even the uncompressed baseline*. Additionally, we show that for both depth upscaling and domain adaptation, a finetuning step of the VAE decoder is sufficient, highlighting the applicability of the data-aware sketching.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.