acceptodds
Under review as a conference paper at ICLR 2027

Who Contributes? Counterfactual Attribution in Sparse Mixture-of-Experts

Abstract

Sparse Mixture-of-Experts (MoE) models have become a key architecture for scaling large language models, yet understanding which experts truly contribute to model performance remains challenging. Existing measures such as routing frequency and gate weight describe expert usage, but overlook that MoE outputs are jointly produced by multiple experts whose contributions may be complementary or redundant. We introduce CONTRAIL, a counterfactual attribution framework that separately measures the contribution of selected expert computation and the utility of router decisions. For expert attribution, CONTRAIL models co-routed experts as a cooperative game and derives closed-form Shapley values from a quadratic approximation of the task loss, capturing both individual and interaction effects with a single backward pass. For router attribution, its Coalition-Consistent Gate Estimator evaluates a mass-conserving counterfactual swap between the nearest selected and excluded experts. We evaluate CONTRAIL on four instruction-tuned sparse MoE models using WikiText-2, GSM8K, MMLU, and MedQA. Across expert removal, cross-task intervention, router analysis, and sparse fine-tuning, CONTRAIL more reliably identifies consequential experts and routing decisions than routing frequency and existing attribution baselines, while revealing task-specific interactions and higher-impact adaptation targets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.