J-MoE: Attention-Aware Joint Sample–Expert Placement for MoE RL Training
Abstract
Reinforcement learning (RL) post-training has become central to improving large language models, which increasingly adopt mixture-of-experts (MoE) architectures. However, variable-length rollout samples and skewed expert routing cause substantial load imbalance across GPUs in three major cost components: dense attention computation, sparse expert computation, and routed All-to-All communication. Balancing these costs requires coordinating two tightly coupled decisions: sample placement and expert placement. Sample placement determines both attention workloads and the sources of routed traffic, while expert placement determines both expert workloads and traffic destinations. Existing load-balancing techniques address only subsets of these costs, without jointly optimizing all three through coordinated sample and expert placement. We present J-MoE, an MoE RL training system that jointly plans sample and expert placement to trade off among all three costs at once. J-MoE leverages routing replay, in which training reuses the routing decisions recorded during rollout, making each sample's expert demand known exactly before training begins. Guided by this information and a unified cost model of all three components, a two-stage solver first assigns the samples of each optimizer step to microbatches and GPUs, and then jointly optimizes sample and expert placement for each microbatch. To execute these plans efficiently, J-MoE redistributes samples to their assigned GPUs and overlaps expert reconfiguration with attention computation, without altering routing or dropping tokens. Across multiple MoE models and RL datasets, J-MoE achieves up to training speedup over existing training systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.