Lateral Inhibition Induces Adaptive Sparsity Resulting in Formation of Interpretable Representations in Autoencoders
Abstract
Lateral inhibition in neurobiological systems achieves sparse activity through recurrent competitive dynamics. Research on sparse autoencoders (SAEs) shows that imposing sparsity facilitates the formation of interpretable features in neural representations. Conventional SAEs, however, typically enforce sparsity through externally specified regularization objectives or activation constraints. To investigate how recurrent lateral competition shapes the learning of interpretable features, we introduce the Lateral-Inhibition Sparse Autoencoder (LISAE), which uses lateral inhibition to generate sparse activity while learning interpretable features in its decoder. Across controlled MNIST datasets and Gemma-2 representations, we show that LISAE recovers interpretable features while enabling input-dependent sparsity, with the number of active neurons adapting to the input structure. Furthermore, we show that anti-Hebbian-trained lateral inhibitory connections capture the relational structure of the feature space, which reflects the underlying similarities among features. Together, these findings establish local lateral inhibition as an alternative mechanism for learning adaptive sparse and relational representations in autoencoders.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.