Autointerpreting Sparse Autoencoders Is a Measurement Problem
Abstract
Autointerpretability scores, where large language models (LLMs) explain and score features, are the predominant way to evaluate sparse autoencoder (SAE) interpretability. Comparing methods with these scores depends on the assumption that they reflect stable feature properties rather than confounding aspects of the evaluation pipeline. Through systematic experiments across four metrics, two models (Pythia-160M, Apertus-8B-Instruct), and six axes of methodological variation, we show that this assumption rarely holds. We find that a single methodological choice explains more score variance than the SAE architecture. Excluding the activation corpus, which is arguably a property of the feature rather than the pipeline, weakens the effect, but it still holds for most model-metric pairs. Moreover, we find that metric reliability is a property of the metric–factor pair, not of the metric, and score agreement does not imply agreement about the top features; a failure that cannot be detected by similar mean scores or highly correlated per-feature scores. These findings suggest that cross-paper comparisons based on autointerpretability scores may reflect pipeline differences rather than architectural differences, providing evidence to an open debate: whether SAE latents are stable, semantically meaningful units for human interpretation. Beyond diagnosing these problems, we propose concrete recommendations for more reliable SAE evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.