acceptodds
Under review as a conference paper at ICLR 2027

GradRouter: Decompose Gradient Before Assigning Learning Responsibility in Sparse Mixture-of-Experts

Abstract

Sparse Mixture-of-Experts (MoE) models enable efficient scaling by activating only a subset of experts for each token. Their effectiveness largely depends on how experts divide and coordinate their responsibilities. Modern MoEs use routers to determine which experts are activated and how much their outputs contribute to the aggregated output. Accordingly, existing algorithms promote expert specialization primarily through forward computation, e.g, differentiating the token distributions and output representations across experts, while leaving backward learning responsibility directly inherited from forward pass. Specifically, for expert outputs are linearly combined, the co-activated experts will receive collinear output-side gradients. In this paper, we raise the attention that such collinear gradient might drive co-activated experts towards the same representation, implicitly hindering their functional differentiation. Moreover, backward learning signals are inherently high-dimensional, with different experts potentially better suited to different directions. This suggests that scalar routing weights alone might not fully characterize how learning responsibility should be divided. Therefore, in this paper, we propose GradRouter, a backward learning responsibility assignment framework that enables differentiated learning direction of co-activated experts. GradRouter first identifies the principal directions of the output-gradient space and then assigns the each direction to co-activated experts according to their capabilities, meanwhile preserving original gradient on aggregated expert outputs. Experiments demonstrate that GradRouter consistently improves model performance across diverse MoE architectures with 7B–30B total parameters, highlighting the effectiveness and scalability of directional learning responsibility assignment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.