MeCeFO-v2: Efficient Fault-Tolerant Optimization for MoE Training
Abstract
Sparse Mixture-of-Experts (MoE) models have become a central architecture for scaling large language models, but they also reshape the fault-tolerance problem during distributed training. Compared with dense models, MoE training couples hardware failures with expert parallelism, routing-dependent workload imbalance, and substantially enlarged training state, making conventional failover strategies inefficient. Existing MoE fault-tolerant methods primarily optimize checkpointing or elastic recovery, but leave open whether the training computation itself can be adapted online after failures with bounded memory and compute overhead. In this work, we propose MeCeFO-v2, a fault-tolerant optimization framework for MoE training under hybrid pipeline, expert, and data parallelism. Instead of assigning a failed worker's workload to a single pipeline neighbor, MeCeFO-v2 redistributes it cooperatively across expert-parallel peers, improving expert update balance and reducing the approximation burden on each surrogate worker. To keep failover training within resource budget, MeCeFO-v2 introduces two structure-aware approximations: norm-guided selective head retention for attention layers and router-aware sequence-level low-rank compression for MoE experts. We further develop an online controller that automatically selects the attention-head budget and expert compression ratio by minimizing profiled gradient error subject to memory and compute constraints. Theoretically, we establish the same convergence rate as standard distributed SGD under bounded gradient error. Experiments on large-scale MoE training demonstrate that MeCeFO-v2 maintains validation perplexity and downstream performance close to fault-free training while limiting throughput degradation to 0.72–6.15%, compared with 7.61–57.92% for prior MoE fault-tolerance baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.