CAER: Cost-Aware Expert Replication for Load Balancing in Mixture-of-Experts Multimodal Large Language Model Training
Abstract
Mixture-of-Experts (MoE) has become a key architecture for scaling multimodal large language models, yet its input-dependent routing can cause load imbalance in expert-parallel (EP) training, leading to stragglers and reduced training throughput. Expert replication is widely used to mitigate such imbalance, but historical load estimates may fail to track rapid workload shifts, while fine-grained, frequent replication introduces additional costs that can offset its benefits. We propose CAER, a cost-aware expert replication method for efficient multimodal MoE (MMoE) training. CAER selectively triggers rebalancing through an adaptive imbalance gate and uses a cost-aware greedy scheduler to identify minimum-cost replication plans under a calibrated cost model that jointly captures computation, communication, and replication overhead. It then dynamically replicates overloaded experts onto underloaded devices while preserving the original routing patterns. Across diverse MMoE architectures and datasets, CAER consistently outperforms state-of-the-art methods, improving end-to-end training throughput by up to 31.8% over EPLB and 31.0% over LPLB while maintaining comparable or better load balance. These results demonstrate the effectiveness and robustness of CAER across diverse training settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.