acceptodds
Under review as a conference paper at ICLR 2027

From Routing Scores to Routing Circuits in Cross-Layer Shared-Expert Transformers

Abstract

Post-hoc circuit tracing can require trained surrogate models whose reconstruction error complicates mechanistic analysis. Building on router-score decompositions and top- selection margins, we study an alternative using cross-layer shared-expert Transformers whose linear routers read the raw residual stream. The difference in linear-router scores between the weakest selected expert and its strongest excluded competitor decomposes exactly over preceding residual writes. This yields a routing-decision graph without a trained explainer. Unlike contributions to individual router scores, these margin contributions are invariant to common shifts of the router weights. We verify the decomposition in models from 4.9M to 81M parameters, with raw-score relative errors of approximately and boundary-margin absolute errors at or below . We then use this accounting to test when routing evidence supports causal conclusions. Routing-related two-dimensional residual subspaces have large effects under ablation, increasing cross-entropy by – nats, but subspaces derived from positive, negative, and total margin contributions largely coincide and have nearly identical effects. Their apparent specificity further weakens under dose-matched controls, and, under the tested interventions, changing expert identity at a single router has little effect on mean cross-entropy. These results establish margin decomposition as an exact, shift-invariant instrument for analyzing expert selection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.