acceptodds
Under review as a conference paper at ICLR 2027

Optimality and Load Balance of Group-Limited Routing in Mixture-of-Experts

Abstract

Mixture-of-experts (MoE) language models activate only a few experts for each token, and large deployments place groups of experts on different devices. Group-limited routing bounds the number of groups per token, but not how much work each group receives. On single-domain batches of real text, the busiest group receives on average nearly twice its fair share; when groups work in parallel, it sets the pace for all of them, and a group with a buffer of fixed size must drop whatever does not fit. We give a novel analysis of group-limited routing that describes each route by how many experts it takes from each group. With this analysis, we determine exactly at which shapes the selection rule released with DeepSeek-V3-style models is score-optimal. We then prove what the existing rules cannot guarantee: with this rule, the routing score on some batches stays well below the best balanced score, whatever the group prices, and on bursty traffic every fixed bias either overloads a group or loses score. Prices set for each batch succeed there. When a linear programme sets one price per group and each token selects exactly its best route at these prices, the programme splits only a few tokens between routes. Rounding them jointly keeps every group within a fixed allowance of its target at any batch size, with a routing score at least that of the best routing that meets the target. On some batches, no routing reaches this score without an allowance. The resulting router, **HYPERSPACE**, needs no retraining and drops nothing. Our empirical results on 159,776 recorded batches of user conversations, including prefill batches of up to 16,384 tokens, show that **HYPERSPACE** meets both guarantees on every batch, whereas the baselines, group prices on the released rule and static replication, miss one of them on some batches. In every one of eight text domains on three MoE models, **HYPERSPACE** has a lower next-token loss than dropping the overflow. On GSM8K and HumanEval, its accuracy stays within four points of the model's own router; dropping loses 5.5 points on GSM8K.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.