Beyond the Previous Layer: Incremental Predictive Value of Routing History in Frozen Sparse Mixture-of-Experts Models
Abstract
Routing in a sparse mixture-of-experts (MoE) transformer is usually examined one layer at a time, yet a single token passes through many routers in sequence. We ask how much older expert selections still predict once the full score vector of the immediately previous router is known. Answering that takes a conditional measurement: we hold that score vector fixed, add older selections, and compare against a distribution-matched control in which history rows are permuted across tokens, which separates alignment from added width. On a frozen grid of four pretrained MoE language models at four depths each, older aligned selections raise held-out for the next router's scores by on average (95% CI , positive in all 16 cells), while the misaligned control recovers none of this gain. The same history raises next-layer expert-set recall from to (), and both effects reproduce on an unseen corpus at four token positions (). The gains vary by nearly an order of magnitude across architectures, and a headroom-normalized diagnostic computed on separate text ranks the cells by realized gain (Spearman ); supplying history to the top half of the cells under that ranking retains about 71% of the full average recall gain. Two conditions bound the scope: widening the conditioning state to three recent score vectors leaves only a small residual, carried by the narrowest router in the grid, and against a strong gate-reading reranker the history-specific complementarity is small and the runtime selector we tested does not convert it. Cross-layer routing history is a real conditional signal whose magnitude is an architectural property rather than a uniform feature of sparse MoE routing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.