acceptodds
Under review as a conference paper at ICLR 2027

HiPipe: Coordinating Shared and Routed Execution in Mixture-of-Experts Models

Abstract

Shared experts in mixture-of-experts (MoE) models process every token to capture common knowledge, complementing the specialization of routed experts. The shared and routed computations can run concurrently, but the layer cannot finish until both outputs are available. In the distributed schedule we study, the shared expert’s first linear projection waits for all input shards, although each shard can be processed independently. We present HiPipe, a hierarchical pipeline that overlaps input transfer with shared projection during routed execution. It projects the local shard first, then overlaps GPU copy-engine transfers of peer shards with projections of ready groups. Configurable grouping balances earlier starts with matrix-multiplication efficiency. Routed work and concurrent execution costs can limit the value of local savings, so we assess groupings by joint completion. We evaluate workloads derived from eleven MoE architectures on H200 and A800. In the repeated A800 study, all eleven shapes improve input projection at 16K and 32K tokens with fixed one-shard groups. At 32K, geometric-mean speedups are 1.416× for input projection and 1.036× for synthetic joint replay over their respective NCCL references. The measured best grouping differs between input projection and joint replay in 27 of 55 workloads. Native-layer and MoE-stack training measurements do not establish corresponding end-to-end gains. These results identify a profitable long-input regime and support choosing grouping and pipeline use by joint completion.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.