Spike-and-Slab Sparse Autoencoders with Adaptive Sparsity
Abstract
Sparse Autoencoders (SAEs) learn sparse representations of deep neural activations, but fixed Top- architectures impose the same feature budget on every input. This aspect limits their ability to adapt representational capacity to inputs of different complexity. In this paper, we introduce Spike-and-Slab Sparse Autoencoders (S-SAE), a Bayesian framework in which the sparsity level can vary using a global learned threshold. By casting sparse encoding as a MAP inference problem, we derive an objective that balances reconstruction fidelity and expected sparsity. We additionally introduce an optional post-hoc refinement that selects among candidate code lengths without retraining. We evaluate our framework on representations from image classification models and Large Language Models, including Pythia-160M and Gemma-2-2B. Results show that S-SAE improves the observed sparsity–reconstruction Pareto frontier across vision and language benchmarks. These findings connect Bayesian sparse coding with practical adaptive sparsity, with potential applications to concept discovery in neural representations. Code is available at https://anonymous.4open.science/r/s2sae.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.