acceptodds
Under review as a conference paper at ICLR 2027

Reinforcing Internal Computations for Large Language Model Reasoning

Abstract

Mixture-of-Experts (MoE) models can solve a problem through many different combinations of experts, yet standard training largely treats expert selection as fixed internal computation. We ask whether reasoning can improve by learning not only what the model outputs, but also how it routes computation internally. We introduce a reinforcement learning approach that jointly optimizes output tokens and expert selection, allowing the model to explore alternative computation paths and reinforce those that lead to better solutions. On Moonlight-16B-A3B, our method improves over token-only GRPO by 2.06 percentage points on mathematical reasoning, 7.49 points on out-of-domain academic reasoning, and 3.83 points on out-of-domain commonsense reasoning. It also reduces average response length by 8.3% under the same training budget, while preserving the standard MoE inference cost. Our results suggest that learning how a model computes can complement learning what it generates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.