Co-Activation Is Not Collaboration: An Interventional Study of Expert Coalitions in Mixture-of-Experts Routing
Abstract
Sparse Mixture-of-Experts routers score each expert independently and execute the Top-, so the routed set maximizes a modular set function, yet everything downstream acts on the selected experts' outputs jointly. Whether set utility is additive cannot be read from co-activation, a statistic of the router's own policy. We measure it interventionally on three frozen MoE language models of 1.3B to 14.3B parameters, swapping coalition members at fixed active compute against a candidate pool that includes uniformly random experts, and decomposing pair utility into unary and interaction terms. Non-additivity is real, shallow and modest in cost. On OLMoE the interaction-energy ratio falls from 0.302 at the first MoE layer to 0.0005 at layer 14, about half the non-additive energy at shallow layers lies beyond any two-slot account, and interaction accounts for about a tenth of the router's regret against the best measured equal-compute coalition; the shallow concentration replicates on all three models. Within-layer pairwise co-activation does not track signed interventional synergy, because a pair's synergy barely reproduces. Co-activation is reliable at the sampling that routing logs supply, but a pair's mean synergy on one set of documents does not predict its mean on a disjoint set ( over eleven layers of two models), and no co-activation–synergy correlation survives FDR in forty-four tests. The oracle gain is heavy-tailed, and a linear per-token head recovers no detectable share of the reweighting gain at our scale, although the same pipeline recovers a third of a planted signal of the same size.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.