Preserving Local Activation Structure for More Reliable Sparse Autoencoder Features
Abstract
Sparse autoencoders (SAEs) provide a way to interpret neural network representations in terms of human-understandable concepts. However, recent evidence of feature instability across training runs and sensitivity to input perturbations raises concerns about the reliability of these interpretations. We argue that preserving the local structure of model activations in their sparse representations allows SAE features to reflect similarities and differences captured by the model. We introduce kNN SAE, which adds a regularizer based on k-nearest neighbors among model activations to incorporate this structural constraint into training. This objective is designed to enable consistent descriptions across related inputs and training runs, making them easier to reproduce and compare. In our experiments, kNN SAE improves feature reproducibility and performance on concept detection and disentanglement evaluations at matched sparsity, with a modest cost in reconstruction and functional fidelity. These findings suggest that local geometric consistency provides a useful foundation for learning more stable and reliable concept representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.