FA-MoE: Fisher-Adaptive Routing for Sparse Vision-Language Mixture-of-Experts
Abstract
Sparse Mixture-of-Experts (MoE) architectures have emerged as an efficient paradigm for scaling large vision-language models (LVLMs). In dense-to-MoE upcycling, experts inherit weights from a pretrained dense model, while the newly introduced router is commonly initialized randomly and trained with an auxiliary load-balancing objective. This leaves two gaps. The first arises at initialization: the experts inherit pretrained structure, whereas the router starts without information about the representation space it must partition. The second arises during training: auxiliary load balancing controls aggregate expert usage but does not directly regulate each token's routing distribution. We introduce Fisher-Adaptive MoE (FA-MoE) to address both gaps. FA-MoE initializes the router with Fisher-discriminative directions derived from the dense model's hidden representations, using K-means cluster assignments as pseudo-labels. During training, it combines token-level routing regularization that (1) controls probability mass outside the selected top-k experts and (2) adapts entropy within the selected set according to routing confidence, with batch-level regularization that maintains balanced expert utilization. Across StableLM-1.6B, Qwen-1.8B, and Phi-2 (2.7B) on eight multimodal benchmarks, FA-MoE ranks first in 21 of 24 backbone-benchmark settings. Controlled ablations show that Fisher initialization performs better than random initialization under the full routing objective, while routing analyses show fewer subsequent changes in the selected expert set.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.