Shaping Latent Mixture-of-Experts: From Expert Geometry to Layerwise Capacity Allocation
Abstract
Fine-grained mixture-of-experts (MoE) models commonly increase expert count by narrowing each expert, treating expert shape as a by-product of routing granularity. LatentMoE moves expert computation into a lower-dimensional space, exposing the input dimension and intermediate width of each expert as explicit design choices. We show that an expert's local Jacobian rank is jointly bounded by these two dimensions, and that balancing them maximizes the rank ceiling at fixed per-expert parameters. Controlled shape sweeps across model scales and routing granularities consistently favor near-square experts. We further use layerwise intrinsic dimension (ID) to guide capacity allocation across depth, co-varying the number and shape of experts while keeping per-layer parameter count, activated FLOPs, and communication volume constant. Compared with layer-homogeneous baselines, our ID-guided allocation achieves lower validation loss and stronger downstream performance. These results establish expert shape and layerwise allocation as practical design axes for MoEs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.