An Identifiability Hierarchy for Mechanistic Interpretability: Interpretability Claims Live at Different Levels of Identifiability
Abstract
Interpretability methods such as sparse autoencoders (SAEs), probes, and steering vectors name units inside a network and claim that those units are useful abstractions. Every such claim presupposes that the unit is identifiable from the activation. We introduce a five-level identifiability framework: three structural levels defined on the code (coefficient, support, and subspace identifiability) and two functional levels defined on the downstream model (behavioral and causal identifiability). We test it on SAEs with one question: is the code the encoder produces the only good sparse code for that activation? Running classical sparse coders on the same dictionary, we find alternatives that overlap with the encoder's code at Jaccard —roughly a third of each code's atoms are absent from the other—while reconstructing the activation as well and the model's predictions slightly better. Support identifiability fails while subspace and behavioral identifiability hold. The codes agree on the atoms with large coefficients and disagree on the many small ones, and the disagreement grows as the SAE is made less sparse, in a regime where the classical conditions for unique sparse recovery do not hold. Coefficient magnitude is the best predictor of how much the model's output changes when an atom is removed; for small atoms it loses significance, and agreement among sparse coders predicts the effect instead. Identifiability, on this view, is not a property an interpretability method has or lacks. It is a level at which a given claim holds—one that a method reaches on a given model and data—and the level at which that claim should be tested.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.