Redesign Mixture-of-Experts Routers with Manifold Power Iteration
Abstract
The router is a cornerstone component of Mixture-of-Experts models. Each row of the router matrix represents an expert, computing its similarity to the MoE input to determine which subset of experts is activated. Ideally, each router row should condense the salient characteristics of its associated expert matrix into a representative vector, such that its dot product can better indicate whether the token should be routed to that expert. However, there is no explicit design principle in existing routers to enforce such a condensation. In this paper, we posit that each router row should be aligned with the principal singular direction of its associated expert matrix, which provides the most expressive one-dimensional representation of the expert matrix. Building on this insight, we propose a router redesign with Manifold Power Iteration (MPI). Specifically, it adopts a “Power-then-Retract” paradigm, in which the power iteration step steers the router weights toward the desired directions, while the retraction enforces a norm constraint to ensure efficient and stable optimization. Our theoretical analysis further confirms that the proposed router design indeed converges toward the target directions. We pretrain MoE models ranging from 1B to 11B parameters and conduct comprehensive experiments to demonstrate the effectiveness of our router design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.