MSLoRA: Multi-Scale Low-Rank Adaptation via Attention Reweighting
Abstract
Parameter-efficient fine-tuning (PEFT) of vision backbones is usually posed as a question of how many parameters to train. We argue that for dense prediction the more consequential question is where a fixed budget should be spent. We study two ways of spending a small budget on a frozen backbone: an additive low-rank residual over intermediate feature maps, and the same residual modulated by an input-dependent, spatially varying multiplicative gate. At budgets of the same order, the multiplicative form is consistently better for localization-heavy tasks and roughly neutral for global classification, which is the pattern one expects if the gate supplies spatial selectivity that a channel-wise low-rank update cannot express. We instantiate this finding as MSLoRA, an adapter that gates a grouped low-rank projection with a multi-scale depth-wise branch. Because MSLoRA acts on feature maps rather than on architecture-specific weight matrices, the same feature-space formulation applies across convolutional, hierarchical-transformer, hybrid, and plain ViT backbones; a separate adapter is still trained per backbone and task, as for other PEFT methods. On COCO with Cascade Mask R-CNN, MSLoRA reaches 42.9 box AP / 38.4 mask AP on ResNet-50 while training 0.7M parameters (2.7% of the backbone), and is competitive with backbone-specialized PEFT methods on Swin-B at roughly half their trainable-parameter count. We report the full cost profile, including the cases where MSLoRA is slower than full fine-tuning, and we audit the strength of the full fine-tuning baselines we compare against.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.