On the Additivity of Expert Contributions in Mixture-of-Experts Models
Abstract
Sparse Mixture-of-Experts (MoE) language models activate a small subset of experts for each token, yet how strongly these experts interact in their effects on the loss remains less well understood. First, we evaluate six open-weight MoE language models on MMLU and compare the loss changes caused by removing experts individually and in pairs within the same layer. We find that removing two experts together changes the loss by approximately the sum of the changes caused by removing each separately, so the effects of experts on the loss are nearly additive. Second, we examine how deviations from this additive relation, which we call interaction, vary with domain overlap and layer depth. Pairs with greater domain overlap exhibit larger absolute interaction when their individual effects on the loss are comparable, while absolute interaction generally decreases in later layers. Finally, we study a single-layer MoE with squared loss, where interaction equals the inner product between expert contributions and vanishes when their outputs are orthogonal. We examine this connection on data generated by orthogonal linear teachers. Training from Gaussian initialization produces near-orthogonal expert outputs and reduces interaction relative to the loss changes caused by removing experts individually. We further prove that training started near the teacher solution recovers it at an exponential rate, yielding orthogonal outputs and vanishing interaction. These results establish this recovery as a sufficient mechanism for additivity in the controlled model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.