HiPipe: Coordinating Shared and Routed Execution in Mixture-of-Experts Models
Abstract
Shared experts in mixture-of-experts (MoE) models process every token to capture common knowledge, complementing the specialization of routed experts. The shared and routed computations can run concurrently, but the layer cannot finish until both outputs are available. In the distributed schedule we study, the shared expert’s first linear projection waits for all input shards, although each shard can be processed independently. We present HiPipe, a hierarchical pipeline that overlaps input transfer with shared projection during routed execution. It projects the local shard first, then overlaps GPU copy-engine transfers of peer shards with projections of ready groups. Configurable grouping balances earlier starts with matrix-multiplication efficiency. Routed work and concurrent execution costs can limit the value of local savings, so we assess groupings by joint completion. We evaluate workloads derived from eleven MoE architectures on H200 and A800. In the repeated A800 study, all eleven shapes improve input projection at 16K and 32K tokens with fixed one-shard groups. At 32K, geometric-mean speedups are 1.416× for input projection and 1.036× for synthetic joint replay over their respective NCCL references. The measured best grouping differs between input projection and joint replay in 27 of 55 workloads. Native-layer and MoE-stack training measurements do not establish corresponding end-to-end gains. These results identify a profitable long-input regime and support choosing grouping and pipeline use by joint completion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.