Why This Expert? A Feature-Level Analysis of Routing in Mixture-of-Experts with Sparse Autoencoders
Abstract
Understanding Mixture-of-Experts (MoE) routing requires identifying which features of a token's representation favor selected experts over their competitors. We present a framework for explaining individual routing decisions through sparse autoencoder (SAE) features. Our framework decomposes routing logits into feature contributions, a decoder bias, and an explicit reconstruction residual, allowing us to quantify how features contribute to routing logits and expert selection. We formulate the search for a minimum-cardinality feature subset preserving the selected expert set as a mixed-integer linear program, distinguishing explanations of reconstructed routing from explanations of original routing conditional on the residual. Across four MoE models, we examine routing fidelity, feature-expert affinities, and the number of features required to preserve selection. On OLMoE, an average of 8.79 of 32 active features suffices to preserve the reconstructed expert set; preserving the original set with the residual fixed requires 9.57 features. SAE reconstructions recover 81–88 of selected experts across OLMoE layers and 73–86 across DeepSeek-V2-Lite layers. Case studies connect these quantitative contributions to interpretable token and contextual features, while steering experiments demonstrate control over expert selection. Together, these results provide a quantitative account of routing through compact feature subsets while explicitly separating their contributions from information the SAE leaves unexplained.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.