Forward-Only LoRA-MoE Allocation via Module Activation Geometry
Abstract
Low-rank adaptation (LoRA) mixtures of experts (LoRA-MoE) increase adaptation capacity through conditional expert activation. However, many conventional LoRA-MoE methods adopted in large language models (LLMs) typically restrict MoE placement to predefined modules, such as the feed-forward networks (FFNs), and use uniform expert configurations across them. This design overlooks that different modules may need different computational capacity for a downstream task. For a more adaptive and more efficient solution, we introduce BLAde (Budget-constrained LoRA Mixture-of-Experts Adaptive Allocation), a task-adaptive framework that jointly determines where MoE capacity should be introduced and how the available factor budget should be allocated across LoRA modules. Using lightweight statistics collected from forward activations, BLAde identifies promising modules for conditional expert specialization, adaptively determines their number of experts and top- routing configurations, and reallocates LoRA ranks across the remaining modules under a factor budget. In this way, BLAde reformulates LoRA-MoE design from a manually specified, architecture-wide configuration into a task-adaptive and budget-aware allocation problem, allocating conditional capacity according to module activation geometry.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.