acceptodds
Under review as a conference paper at ICLR 2027

Expert-Space Exploration in MoE Reinforcement Learning

Abstract

Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency. Since routing determines the sparse computation paths that induce output distributions, expert selection offers an additional source of rollout diversity beyond conventional token-level sampling. Through empirical analysis, we find that routing perturbation changes model outputs and increases rollout diversity, whose effect resembles that of increasing the decoding temperature. However, direct perturbation can activate unsuitable experts and substantially degrade rollout quality. Motivated by these observations, we introduce **E**xpert-**S**pace Exploration **R**einforcement **L**earning (**ESRL**), an architecture-aware framework that explicitly explores the expert-routing space of MoE models. ESRL preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths. The perturbation strength is further adapted according to router entropy to avoid over-perturbation. To mitigate the routing mismatch introduced by perturbation, ESRL records the expert paths used during rollout and replays them during policy optimization. Experiments demonstrate that ESRL achieves the best performance across MoE backbones with top-, top-1, and shared-expert routing on mathematics, science, and code tasks, without increasing the number of rollout samples or activated experts per token. Specifically, ESRL on Qwen3-30B-A3B improves average Pass@1 and Pass@8 over GRPO by **3.2** and **4.5** percentage points, respectively. Further analyses of expert utilization and training dynamics provide insights into how exploiting MoE-specific routing structure benefits RL training. These results establish expert routing as an effective and complementary dimension for RL exploration beyond its architectural role.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.