acceptodds
Under review as a conference paper at ICLR 2027

A Unified Theoretical Framework of Sparse Dictionary Learning in Mechanistic Interpretability

Abstract

As AI models achieve remarkable capabilities across diverse domains, understanding what representations they learn and how they encode concepts has become increasingly important for both scientific progress and trustworthy deployment. Recent works in mechanistic interpretability have widely reported that neural networks represent meaningful concepts as linear directions in their representation spaces and often encode diverse concepts in superposition. Various sparse dictionary learning (SDL) methods, including sparse autoencoders, transcoders, and crosscoders, are utilized to address this by training auxiliary models with sparsity constraints to disentangle these superposed concepts into monosemantic features. These methods are the backbone of modern mechanistic interpretability, yet in practice they consistently produce polysemantic features, feature absorption, and dead neurons, with very limited theoretical understanding of why these phenomena occur. Existing theoretical work is limited to tied-weight sparse autoencoders, leaving the broader family of SDL methods without formal grounding. We develop the first unified theoretical framework that formulates major SDL variants within a common optimization problem and show that, under sparse linear representation assumptions, its feature-wise approximation exhibits piecewise biconvex structure. We characterize the zero-loss conditions and spurious partial minima of this objective, exposing the underdetermined nature of reconstruction-based SDL and providing principled explanations for feature absorption and dead neurons. Furthermore, we demonstrate how to reduce this underdetermination with a novel technique, feature anchoring, as an extension of our theory.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.