WAS-LoRA: Output-Sensitive Subspace Selection for Post-Training LoRA Compression
Abstract
We study post-training LoRA compression to reuse trained adapters under tighter rank budgets at deployment, without requiring an additional fine-tuning stage. Small reconstruction errors in adapter weights or the local LoRA branch outputs, however, do not necessarily ensure that the compressed model's final logits remain close to those of the original adapted model. We propose WAS-LoRA (Whitened Active Subspace Low-Rank Adaptation), which guides active-subspace selection using the sensitivity of a selected teacher output downstream of the adapter's bottleneck. It constructs a Jacobian Gram matrix that measures the output sensitivity in the adapter's whitened bottleneck coordinates. The leading eigenvectors of this Jacobian Gram matrix span the retained subspace, from which we directly construct standard lower-rank LoRA factors. We prove that, for a fixed retained rank, this subspace minimizes a Jacobian-based measure of output sensitivity along discarded bottleneck directions. Experiments on GLUE classification tasks across five different rank budgets demonstrate that WAS-LoRA reduces geometric-mean centered relative logit error by 30.1% and 15.8% compared with the activation-aware and fixed-basis baselines, respectively. Task-accuracy differences among methods are concentrated at the tightest rank budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.