acceptodds
Under review as a conference paper at ICLR 2027

Auditing Expert Architecture Regret in Mixture-of-Experts Language Models

Abstract

Sparse Mixture-of-Experts models fix their expert architecture before training, while adaptive methods reorganize it during training when an observational statistic fires. Both rest on one unmeasured premise, that the changes such statistics suggest actually improve the model. We measure it. Expert Architecture Regret is the largest validation-loss improvement available from an audited local operation (retire, merge, or a budget-neutral composite), estimated by paired interventional probes against a null of no-op pairs. Each probe adapts two branches identically except for the operation. Across 407 probe results spanning two fine-grained families, seven checkpoints, random and heuristic candidates and two data distributions, six candidates cross the band. None survives re-probing on all control streams, where function-preserving placebos score the same. The instrument is not blind, since an amplitude control puts a known defect above the band and the audit finds it. But its resolution turned out to be a protocol choice. The adaptation rate moves the un-mutated model ten to fifty times further than any single expert can. Lowering it tenfold cuts the band from 0.0041 to 0.0025 nats while the control still fires, and 101 further operations give a largest effect of +0.0003. A sixteen-stream re-test confirms two of them, near-dead-expert retires, as gains of 1–2×10⁻⁴ nats that the pooled rule cannot flag. These gains hold on held-out WikiText text and reverse on OLMoE's own training mixture, where the two experts are not near-dead. Destroying a layer's busiest expert costs 0.0018–0.0100, so a removal gain of that size reaches 80% power only where experts are large. We state this scale argument and give a counterexample class to it. A coarse-grained model (JetMoE-8B, 8 experts, top-2) confirms the account. In it, a layer's busiest expert is worth 11–26× its band, and eight of nine audited edits lose, six of them by 2.5–13× that band. The ninth, which only a random candidate proposed, retires a busiest expert that is harmful on WikiText. It clears the band on all six streams of the validation slice, vanishes on the held-out test split (−0.0001, 95% CI −0.0004 to +0.0002) and reverses on RefinedWeb, JetMoE's main training source (−0.0036, every stream negative). Every positive audited LM edit we find reverses on data its model was trained on. Half a null draw's variance separates repeat runs of one configuration rather than one stream from another. More streams shrink it and 16× the evaluation tokens do not, so buy control streams, not evaluation data. We release the harness, the raw per-run records, and the scripts that recompute the statistics quoted here.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.