Learning Balanced Sparse Routing with Uniform Pairwise Distillation
Abstract
Sparse Mixture‑of‑Experts (MoE) layers activate a small subset of experts per token, so practical runtime efficiency heavily depends on how evenly the router distributes token‑level demand. Skewed demand creates overloaded hot experts, underutilized experts, rank imbalance, and extra runtime repair overhead. Existing approaches either apply indirect global balancing losses, or only perform ephemeral runtime corrections. We introduce Uniform Pairwise Distillation (UPD), a lightweight local self‑distillation learning objective for training better‑balanced router proposals. For every route slot modified by the strict capacity‑aware execution policy during training, UPD constructs one sparse pairwise ranking target that favours the accepted replacement expert over the displaced expert. It assigns equal weight to each observed correction pair, reuses the router’s own logits as student signals, and introduces neither an external teacher network nor additional trainable parameters. Across matched from‑scratch MoE configurations ranging from 350M to 7B total parameters, UPD reduces raw‑demand imbalance and runtime reassignment work while maintaining perplexity comparable to the no‑distillation baseline. Improvements in routing behaviour transfer to held‑out in‑domain and out‑of‑domain test inputs. In a two‑rank expert‑parallel pipeline, these learned routing gains translate to higher prefill throughput and modest, repeatable decode improvements, while the original execution policy continues to enforce hard capacity feasibility guarantees.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.