: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
Abstract
Mixture of Experts (MoE) Large Language Models (LLMs) achieve strong performance at scale. However, reinforcement learning (RL) on MoE-based LLMs often suffers from training instability. A root cause is router drift, i.e., expert activations can change drastically across model updates and differ between disaggregated rollout and training phases, causing large rollout–training mismatch and unstable importance sampling weights in PPO-style RL algorithms. Routing replay mitigates this issue by freezing the replay route within each reasoning trajectory, but it ignores how the router evolves under off-policy updates and thus causes router staleness. We introduce Predictive Routing Replay (PR2), which augments each router with a lightweight evolution predictor that learns to anticipate short-horizon router evolution. During route recording, the predictor selects expert sets that anticipate router evolution. During training, it replays these sets with the current model. We bound recorded-route sensitivity through predictor fit and update drift, derive an exact KL threshold for top- agreement, and show how local support certificates yield matching forward computations and fixed-batch policy gradients. Across three MoE models and reuse factors , and , PR2 achieves the highest average accuracy in all nine comparisons, improving average math accuracy over routing replay by to percentage points on Qwen3-30B-A3B-Base.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.