MoE Routing Is Not Token-Local: How Expert Influence Propagates Across Tokens & Layers
Abstract
In Sparse Mixture-of-Experts (MoE) models, routing is typically analyzed at the level of individual tokens. We instead trace how an expert's computation on one token affects the routing decisions of other tokens. We show that routing influence propagates across tokens and layers through two causal channels, a direct same-token channel and a context-mediated inter-token channel. Our evaluation on five MoE models reveals that: (i) the two channels combine approximately linearly in routing-margin space before the first routing divergence; (ii) they propagate differently with depth, as direct influence acts immediately and fades while context-mediated influence accumulates downstream; and (iii) the two channels are typically redundant rather than synergistic, with either pathway alone often sufficient to cross a routing boundary. To formalize these findings, we examine MoE routing boundaries and introduce a first-order, margin-normalized routing leverage. Routing leverage makes explicit that whether a Top-K boundary is crossed depends on how an expert's contribution aligns with the contrast direction between the competing experts, relative to the margin separating them, and it predicts the layer at which a token's routing first diverges, with accuracy increasing as more of the causal path is included. Discrete Top-K selection turns these continuous, additive effects into nonlinear changes in computation, and the first divergence marks the onset of a routing cascade.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.