acceptodds
Under review as a conference paper at ICLR 2027

Sparsely gated tiny linear experts

Abstract

Sparsity allows scaling model parameters without proportionally increasing the size of active circuits. While mixture of experts models are made increasingly sparse, individual experts typically remain large and dense. Here, we demonstrate that further increasing sparsity by shrinking each expert to consist of a single neuron and selecting a tiny fraction of many available neurons can improve mechanistic interpretability and compute-matched performance. Counterintuitively, the key to achieving both is removing the nonlinearity typically applied to the experts, resulting in a network of sparsely gated linear neurons (sgatlin). In an isoflop comparison, we find that replacing all transformer feedforward layers with sgatlin improves perplexity in language models across different compute budgets. Crucially, the sparsity and linearity of the resulting feedforward circuits simplifies mechanistic analyses. In a small-scale case study, we demonstrate that the sparse feedforward circuits in sgatlin can be interpreted without having to train additional replacement models (e.g., sparse autoencoders). We find that feedforward circuits live in a semantically structured metric space and can be selectively intervened on to probe their causal roles. Our findings paint a possible path towards interpretable transformers without sacrificing modelling quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.