acceptodds
Under review as a conference paper at ICLR 2027

MoE Routing Is Not Token-Local: How Expert Influence Propagates Across Tokens & Layers

Abstract

In Sparse Mixture-of-Experts (MoE) models, routing is typically analyzed at the level of individual tokens. We instead trace how an expert's computation on one token affects the routing decisions of other tokens. We show that routing influence propagates across tokens and layers through two causal channels, a direct same-token channel and a context-mediated inter-token channel. Our evaluation on five MoE models reveals that: (i) the two channels combine approximately linearly in routing-margin space before the first routing divergence; (ii) they propagate differently with depth, as direct influence acts immediately and fades while context-mediated influence accumulates downstream; and (iii) the two channels are typically redundant rather than synergistic, with either pathway alone often sufficient to cross a routing boundary. To formalize these findings, we examine MoE routing boundaries and introduce a first-order, margin-normalized routing leverage. Routing leverage makes explicit that whether a Top-K boundary is crossed depends on how an expert's contribution aligns with the contrast direction between the competing experts, relative to the margin separating them, and it predicts the layer at which a token's routing first diverges, with accuracy increasing as more of the causal path is included. Discrete Top-K selection turns these continuous, additive effects into nonlinear changes in computation, and the first divergence marks the onset of a routing cascade.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.