acceptodds
Under review as a conference paper at ICLR 2027

Preserving Local Activation Structure for More Reliable Sparse Autoencoder Features

Abstract

Sparse autoencoders (SAEs) provide a way to interpret neural network representations in terms of human-understandable concepts. However, recent evidence of feature instability across training runs and sensitivity to input perturbations raises concerns about the reliability of these interpretations. We argue that preserving the local structure of model activations in their sparse representations allows SAE features to reflect similarities and differences captured by the model. We introduce kNN SAE, which adds a regularizer based on k-nearest neighbors among model activations to incorporate this structural constraint into training. This objective is designed to enable consistent descriptions across related inputs and training runs, making them easier to reproduce and compare. In our experiments, kNN SAE improves feature reproducibility and performance on concept detection and disentanglement evaluations at matched sparsity, with a modest cost in reconstruction and functional fidelity. These findings suggest that local geometric consistency provides a useful foundation for learning more stable and reliable concept representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.