acceptodds
Under review as a conference paper at ICLR 2027

Sparse Autoencoders Do Not Guarantee Local Features

Abstract

Sparse autoencoders (SAEs) can recover sparse feature coordinates without identifying which inputs determine their activations. A feature associated with one input pattern may still depend on several controlled variables, and objectives based only on representation samples cannot distinguish input maps that induce the same representation distribution. We introduce the optimal input scope curve, which measures the least activation variance left unexplained when feature scopes are assigned under a shared dependency budget. For every interaction order , we construct systems with identical representation distributions and the same frozen SAE whose localization budgets differ by exactly a factor of at the same omission tolerance. This separation persists when an exact sparse coding objective uniquely selects the shared dictionary. We further derive compatibility and spectral lower bounds linking joint input effects, decoder directions, and required scope budgets, together with finite sample guarantees for paired intervention estimates after scope selection. In five learned overcomplete SAEs, two bijective input maps increase the mean budget needed to explain 90% of activation variance by 41% and 49%, despite leaving the representation and activation samples unchanged. These results show that sparsity and feature recovery do not by themselves establish which inputs determine a sparse feature.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.