When Can Compression Improve LoRA? A Bias–Variance Theory of Restricted Adaptation
Abstract
Compressed variants of LoRA can sometimes outperform standard LoRA despite using far fewer trainable parameters. We study this effect as restricted estimation: compression reduces estimation variance but introduces task-dependent projection bias. A local quadratic decomposition and an exact sequence model yield an energy–noise condition for when restriction helps. At matched parameter budgets, sparse residual correction is beneficial when its joint recovery exceeds the opportunity cost of displaced compressed capacity, with an additional retained-noise term under adaptive selection. For Uni-LoRA-style parameter sharing, we characterize recovery after selected coordinates are released and the remaining shared parameters are refitted. This motivates Projected LoRA with Sparse Adapter (ProLoSA), which combines a compressed branch with selected residual coordinates. Sequence and Transformer controls test the mechanism and its limits, while downstream experiments on language understanding, reasoning, instruction tuning, and vision show matched-budget gains over compressed LoRA in several settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.