FactorEP: Accelerating Expert-Parallel MoE Inference through End-to-End Factor-Space Execution
Abstract
Mixture-of-Experts (MoE) models rely on expert parallelism (EP) to scale across GPUs, but communication can account for up to 47% of end-to-end execution time. Existing EP systems primarily optimize how activations are communicated, while redundancy in pretrained experts offers an opportunity to reduce what must be communicated and computed. We present FactorEP, a system that exploits this redundancy by co-designing compact representations and EP execution. Group-wise Tucker factorization with KFAC-weighted calibration produces group-shared input factors, expert-specific compact cores, and a layer-shared output factor. This asymmetric sharing enables input projection before dispatch and output reconstruction after combine, reducing both communication and expert computation. To turn these reductions into practical speedups, a three-stage factor-space pipeline amortizes shared projection costs and overlaps communication with compact expert execution. Route-compiled execution uses a packed token layout shared by communication and compact expert kernels to reduce capacity padding and per-expert launch overhead. On Qwen3-30B-A3B, FactorEP achieves higher average downstream accuracy than all evaluated compression baselines across all tested parameter budgets. Its deployment configuration reduces dispatch and combine feature widths by 20% while maintaining average downstream accuracy within 1.93 percentage points of the original model. The same configuration achieves end-to-end speedups over COMET of on eight H100 GPUs with NVSwitch and on eight RTX 6000 Ada GPUs over PCIe. These results, together with quality-latency sweeps on Qwen3-30B-A3B and OLMoE-1B-7B, demonstrate FactorEP's applicability across MoE models and hardware platforms.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.