acceptodds
Under review as a conference paper at ICLR 2027

Learning Balanced Sparse Routing with Uniform Pairwise Distillation

Abstract

Sparse Mixture‑of‑Experts (MoE) layers activate a small subset of experts per token, so practical runtime efficiency heavily depends on how evenly the router distributes token‑level demand. Skewed demand creates overloaded hot experts, underutilized experts, rank imbalance, and extra runtime repair overhead. Existing approaches either apply indirect global balancing losses, or only perform ephemeral runtime corrections. We introduce Uniform Pairwise Distillation (UPD), a lightweight local self‑distillation learning objective for training better‑balanced router proposals. For every route slot modified by the strict capacity‑aware execution policy during training, UPD constructs one sparse pairwise ranking target that favours the accepted replacement expert over the displaced expert. It assigns equal weight to each observed correction pair, reuses the router’s own logits as student signals, and introduces neither an external teacher network nor additional trainable parameters. Across matched from‑scratch MoE configurations ranging from 350M to 7B total parameters, UPD reduces raw‑demand imbalance and runtime reassignment work while maintaining perplexity comparable to the no‑distillation baseline. Improvements in routing behaviour transfer to held‑out in‑domain and out‑of‑domain test inputs. In a two‑rank expert‑parallel pipeline, these learned routing gains translate to higher prefill throughput and modest, repeatable decode improvements, while the original execution policy continues to enforce hard capacity feasibility guarantees.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.