Reinforcing Internal Computations for Large Language Model Reasoning
Abstract
Mixture-of-Experts (MoE) models can solve a problem through many different combinations of experts, yet standard training largely treats expert selection as fixed internal computation. We ask whether reasoning can improve by learning not only what the model outputs, but also how it routes computation internally. We introduce a reinforcement learning approach that jointly optimizes output tokens and expert selection, allowing the model to explore alternative computation paths and reinforce those that lead to better solutions. On Moonlight-16B-A3B, our method improves over token-only GRPO by 2.06 percentage points on mathematical reasoning, 7.49 points on out-of-domain academic reasoning, and 3.83 points on out-of-domain commonsense reasoning. It also reduces average response length by 8.3% under the same training budget, while preserving the standard MoE inference cost. Our results suggest that learning how a model computes can complement learning what it generates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.