MoE Mirage: Selection Noise and Routing-Weight Geometry in Token-Level Counterfactuals
Abstract
Token-level "best expert" headroom in sparse mixture-of-experts (MoE) models is a same-sample best-of- maximum: each candidate expert is installed, the downstream loss is scored, and the best is kept, on the same tokens. We show that, at the resolution we can test, this quantity behaves as selection noise whose scale is set by the intervention operator. Selected on one half of the continuation and scored on the other, no operator retains headroom above the experiment's resolution (identity-preserving swap: nats, 95% CI ; planted advantages detected with 80% power from ), and an exact matched-noise null closely reproduces the primary same-sample values and the checkpoint ordering of their ratio. Published operators change support, weight shape, total mass or normalization together with the expert; we introduce an identity-preserving swap that changes only the expert, and full-commitment forcing reports more same-sample headroom than it in of cells (), as the null predicts. The operator also changes a decision that takes no maximum: one-shot pruning by mean expert effect yields operator-dependent masks, forcing's ranking partly reproduces across document halves while the swap's does not, and on two OLMoE releases forcing-ranked masks fall below random masks in benchmark accuracy. We recommend that token-level routing maps state their operator and scoring horizon, cross-fit candidate selection and scoring, and report a matched-noise null to separate persistent opportunity from selection noise.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.