acceptodds
Under review as a conference paper at ICLR 2027

MorphSAE: Reserved Sparse Slots without Automatic Component Recovery

Abstract

Understanding a neural network may require separating its activations into components with different roles. Standard TopK sparse autoencoders learn one sparse feature inventory, so they cannot ensure that every proposed group receives selected features. We introduce MorphSAE, an additive sparse autoencoder inspired by Morphological Component Analysis (MCA), a sparse-demixing method from the Compressed Sensing and sparse-representation literature. Each named dictionary has its own TopK budget, and the decoded components add to reconstruct the original activation. This design guarantees returned indices from every dictionary. In three matched comparisons, an optional cross-dictionary loss also reduced its sampled overlap statistic. In a controlled synthetic copy task, an explicit separator signal concentrates the task-relevant activation change in a controller dictionary. Compared with a selected conventional-SAE intervention, the combined named and supervised design improves target recovery while reducing collateral error. This demonstrates a useful named decomposition in this supervised synthetic task. A matched model with one shared budget performs similarly, so the quota alone does not explain the success. On natural activations, the tested quota provides no reliable reconstruction or unsupervised-specialization benefit. Because each decoded component remains in the original activation space, the architecture permits it to be passed into another MorphSAE. We hypothesize that future separator functions could support repeated decomposition into smaller, more interpretable dictionaries; this remains untested.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.