acceptodds
Under review as a conference paper at ICLR 2027

Batch-Aware Hierarchical Routing for Mixture-of-Experts with Reduced Memory Traffic and Theoretical Guarantees

Abstract

Mixture-of-Experts (MoE) architectures reduce computation by activating a small subset of experts per token. However, under batch-based execution, memory access is determined by the union of experts selected across the batch, so memory movement remains a key bottleneck. We propose Batch-Aware Hierarchical MoE Routing (BAR), which augments token-level routing with a batch-level router that first selects a small subset of experts per batch, followed by token-level refinement. This design explicitly reduces memory traffic while improving efficiency and expert specialization. We theoretically show that BAR reduces memory access, inference cost, and training computation, established through the training dynamics analysis of two-layer ReLU experts. Across evaluation on large-scale MoE language models, BAR substantially reduces HBM traffic and improves inference throughput while closely preserving base-model performance. For example, on DeepSeek-MoE-16B, BAR reduces HBM traffic by 60% and improves throughput by \(2.5\times\), while maintaining perplexity within 2% of the base model

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.