acceptodds
Under review as a conference paper at ICLR 2027

When a Load-Balancing Loss Rewards Dead Experts: Sharp Global Geometry of Switch Routing

Abstract

Top-1 mixture-of-experts routers are commonly regularized by the Switch auxiliary loss , which couples hard expert frequencies with mean soft routing probabilities . It is widely believed that the loss attains its minimum value of 1 under uniform routing. We show that this claim is false. We prove the sharp infimum The bound is approached by routing all tokens evenly among roughly *half* the experts, leaving the rest unused. For each token, the routing probabilities approach a uniform distribution over its assigned expert and all unused experts, with the assigned expert remaining the unique top-1 choice. We characterize all extremal hard loads and prove that hard-load vectors must approach this half-dead family as the loss approaches its infimum. We also construct cases where the loss decreases while expert assignments and routed outputs remain unchanged under top- routing with post-selection renormalization. We explicitly observe normalized losses below 1 in layers of released Switch Transformer, Qwen, and DeepSeek models. In Switch Transformer and Qwen, such values coexist with unused experts that retain above-uniform soft mass. Controlled language-model training experiments also produce losses below 1 alongside unused experts and show that Switch regularization can induce *negative* hard–soft correlation. These results show that minimizing the auxiliary loss does not necessarily balance expert utilization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.