Same Features, Opposite Conclusions: Concept Assignment in SAE Evaluation for Mechanistic Interpretability
Abstract
Sparse autoencoders (SAEs) help researchers study the features learned by neural networks, but only a small subset of these features can usually be inspected. These features may correspond to concepts which are human-interpretable properties or categories. Choosing that subset of features requires a way to score features and evaluate the selections of concepts. We investigated whether the relative performance of two feature-selection methods (top-480 and uniform-480 varies depending on the concepts used for evaluation. Across 163 trained SAE checkpoints, we compared feature scores calculated using either the 480 most strongly activating examples or 480 randomly sampled active examples. We represented each feature using its decoder direction and measured how well it separated active from inactive examples. We assigned each selected feature the concept that maximized either binary F1 or absolute cosine similarity between its decoder direction and a known concept direction. We then evaluated concept detection on held-out examples using precision, recall, and F1 between feature and concept presence. Across matching-pursuit and standard SAEs, changing the concept assignment rule reversed which selection method achieved higher mean F1, even though the feature scores and selections were unchanged. The reversals recurred with fresh scoring, assignment, and evaluation data and new candidate pools, although statistical support was weaker for matching-pursuit SAEs under a common-seed analysis. The original test across 105 held-out checkpoints did not support the hypothesis that uniform sampling selects better concept detectors overall. In a separate experiment, training SAEs on data where concepts occurred independently greatly reduced disagreement between the assignment rules, even when the models were evaluated on the same data. These results show that comparisons between feature-selection methods can depend on how concepts are assigned to features. Although our experiments do not establish generalization beyond one synthetic generator or identify a universally better method, they show that changing the evaluation target can reverse conclusions while feature scores and selections remain fixed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.