Same Counts, Different Cost: Measuring the Order Gap in Sparse Mixture-of-Experts Routing
Abstract
Adapt a sparse Mixture-of-Experts model to a new task and you can easily measure how much its routing changed: how often each expert fires, how far the choices have moved, which experts fire together. Those summaries are then used to decide what to keep resident in memory. Every one is a count, a rate or a set overlap, and none records the order in which experts are accessed. We call the cost of that omission a trace's order gap: what randomizing token order does to simulated miss rate while every one of the trace's own counts is preserved exactly. With 16 of 64 experts resident per layer it is 0.135 on full traces, holding to within 0.005 across two checkpoints, an independently adapted second run and a sixfold increase in coverage, while adaptation itself moves the same metric by 0.0008. A separate comparison at twice the optimizer budget, on its own population, shifts the mean order gap by 0.000232. Replayed on real expert-sized tensors on an RTX A6000, randomizing token order adds 14.2 seconds of expert transfer to the actual order's 38.9, a 36.6% increase. Order gaps also vary across real checkpoints, by more than adaptation moves the metric over 225 checkpoint-layer cells from four adapted runs; each carries its own count-matched baseline, so that comparison is descriptive. The insufficiency is then exact rather than comparative. A constrained kernel additionally holds each prompt's aggregate overlap with the pre-adaptation checkpoint fixed, so per-prompt expert counts, per-expert rank counts and aggregate overlap are identical on both sides, and the configuration it reaches over 120 prompts and all 15 eligible layers, invariants checked on each of the 1,800 cells, differs by +0.015465 in miss@16; replayed on a 20-prompt, five-layer subset it adds 2.25 seconds of expert-loading time. One such pair shows that no function of those summaries alone reproduces the simulated cost; our sampler does not mix, so those figures describe the configuration reached rather than a population effect. Behind the constructions is the measurement program that calibrates them. Under an empirical null stratified by task, router rank, resident-set size and token position, newly selected experts show excess recent-window residency of +0.037129 and +0.037296 in two OLMoE-1B-7B trajectories and three to six times less in Qwen1.5-MoE-A2.7B, while matched comparisons give no evidence that the longitudinal residual exceeds that between two independently adapted runs. The claim is narrow: these summaries do not determine the cost they are used to reason about, and the gap is worth a third of this workload's expert transfer time on real hardware. The replay measures transfer in isolation, so we claim no end-to-end benefit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.