Feature Sparsity via Margin Maximization
Abstract
Explaining the number of features in a network has been a long-desired result in mechanistic interpretability. [Morwani et al. (2024)](https://arxiv.org/abs/2311.07568v2) laid the baseline through their margin-maximization analysis, proving that quadratic networks maximizing margin under matched regularization use every frequency for modular addition. This restriction, however, excludes most models conventionally trained using regularization or a non-quadratic architecture, which are empirically sparse. Our inquiry focuses specifically on modular addition in polynomial and multilinear models. Our primary theoretical finding centers on the cost of new frequencies—a trade-off between the margin of the solution and regularization. For interference-free product networks, we use this to characterize the maximum normalized margin and an optimal number of frequencies for a particular configuration of homogeneity degree () and regularization. This recovers the dense frequency allocation seen in [Morwani et al. (2024)](https://arxiv.org/abs/2311.07568v2). Beyond the interference-free class of models, we also prove upper and lower bounds on the maximum normalized margin of product networks. We also establish matching bounds on the optimal number of frequencies, and prove that an optimal spectral allocation uses frequencies. Finally, we prove a general asymptotic scaling law for normalized margin over product models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.