acceptodds
Under review as a conference paper at ICLR 2027

EquiMoE: A Mixture-of-Experts Architecture with Load Balance by Construction

Abstract

Sparse mixture-of-experts (MoE) routing can concentrate assignments on a few expert groups, creating uneven work across expert-parallel ranks. EquiMoE fixes each group's assignment budget in the forward operator: every token enters every learned subspace branch, where an independent router selects a fixed number of experts from a disjoint bank. This gives exactly equal group assignment counts for any input and router parameters, while expert loads within each group may remain unequal. With one branch per rank, EquiMoE sends one reduced representation per token to each rank before local routing and returns one aggregated result per branch. The nominal per-call all-to-all payload is 4.19 MB, versus 33.55 MB for an assignment-expanded Standard MoE baseline: an 87.5% reduction under this dispatch convention. In four C4 runs of 20,000 steps (2.62 billion processed token positions), EquiMoE has equal rank assignment counts, constant measured payload, and training-loss trajectories close to the baselines. A LatentMoE-style scale-up comparison shows that reduced expert width with assignment-expanded dispatch does not yield the same payload saving. The guarantee concerns group assignments, not equal per-expert load. While EquiMoE's per-call communication time is approximately 26% lower than the baseline, this does not translate to proportional end-to-end speedup due to fixed all-to-all latency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.