SVD-IMOE:PRETRAINEDSPECTRALDIRECTIONSAS IMPLICITROUTERSFORLOW-RANKADAPTATION
Abstract
Parameter-efficient fine-tuning with low-rank adapters typically relies on dense rank usage or randomly initialized routing directions, which may ignore the spectral geometry of pretrained weights. We propose SVD-IMoE, a spectral implicit mixture-of-experts framework that uses pretrained right-singular directions as frozen routing bases for sparse low-rank adaptation. SVD-IMoE-W selects directions through calibration-free stratified Weight-SVD and applies inverse-square-root spectral scaling. SVD-IMoE-TA further uses task activations to select task-aligned directions within the same spectral bins, together with a heuristic alignment gate and damped adaptive normalization. Both variants use token-wise top- routing and train only the LoRA up-projection. With matched and , SVD-IMoE-W improves GSM8K over FlyLoRA by 4.39 points on Llama 3.1-8B and 2.27 points on Qwen 2.5-7B. SVD-IMoE-TA improves Qwen HumanEval Pass@1, Pass@5, and Pass@10 over SVD-IMoE-W by 3.60, 3.27, and 2.21 points, respectively. In four-adapter parameter averaging, SVD-IMoE-TA outperforms merged FlyLoRA on all twelve evaluated backbone-metric comparisons. These results suggest that pretrained spectral geometry provides a useful routing prior, while activation geometry enables selective task specialization and improves cross-task adapter compatibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.