PrismLoRA: Learning Rate as Structure in Mixtures of LoRA Experts
Abstract
Mixtures of LoRA experts (MoE-LoRA) are reported to outperform LoRA in multi-task learning. However, they outperform LoRA at a shared learning rate, not at LoRA's finely tuned one. We show that MoE-LoRA structures mainly act as hidden learning-rate multipliers, i.e., finely tuned learning rates remove 89% (vision) and 96% (language) of the accuracy difference between methods. We further show that tasks need different learning rates, which one shared multiplier cannot provide. For instance, on the same smallNORB images, the camera's azimuth needs a large learning rate and its elevation a small one. We explain these differences with a Bayesian update, where the learning rate sets the gain (i.e., how far training moves the pretrained model toward the labels). Because a confident prior needs a small gain, we measure this confidence with frozen features and predict before training which tasks need a small learning rate. We confirm every prediction after training. We also show that tasks should share parameters only when they need the same learning rate. We therefore propose PrismLoRA, which splits the LoRA rank into blocks at different learning rates and lets each task use the blocks near its optimal learning rate. PrismLoRA sets each task's blocks automatically from frozen features and validation curves. PrismLoRA outperforms all eight tuned baselines with matched parameters. Specifically, PrismLoRA improves over the strongest baseline by 2.63 points on VTAB-6 and 1.35 on Structured-8, and ranks first on language tasks. How much to trust the prior should therefore be set per task by the model's structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.