Why Better Expert Choices Do Not Always Make Better MoE Routers
Abstract
Sparse mixture-of-experts (MoE) routing determines which experts contribute to each next-token prediction in language models. A router selects a small set of experts from the available prefix before the next token is observed. Yet a search conducted after observing that token can find expert choices that predict it much better than the native route, leaving unclear how much of this gain a router could achieve. Recent methods use such hindsight choices to supervise router training. Because the search sees the answer that the router cannot, part of its gain may be unavailable to any input-only decision. Conversely, a trained router's failure to match the search cannot show whether it has also missed useful choices that were available before the answer. We separate these effects within fixed candidate sets by using independent human completions of the same prefix to lower-bound the value of answer information without assuming an optimal learner. We then use a frozen reference model's complete answer distribution to compute both the hindsight and best input-only values and test whether trained selectors recover the latter on new prefixes. Across three frozen MoEs, answer information accounts for a provable part of the search gain, while the tested selectors leave input-only value unrecovered; in the controlled reference setting, about 70% of the search gain requires the answer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.