When AI Explains AI: Understanding and Mitigating the Interpretability Lottery
Abstract
Interpreting the internal representations of language models is crucial for understanding their behavior and guiding efforts to make them reliable and safe. Recent advances have enabled the discovery and interpretation of model features at scale. We study the interpretability lottery, in which a feature’s attributed meaning remains contingent on the evidence and inferential choices of the interpretation procedure after evaluation and selection. A single published description leaves this semantic uncertainty implicit. We ask how the uncertainty relates to a feature’s activating evidence and use, and how it can be mitigated. First, we probe an explain-then-score pipeline on Gemma Scope sparse autoencoder (SAE) features of Gemma-2-2B, rerunning it fifteen times on each of 150 features with the model, the SAE and the feature held fixed. No feature is near-unanimous across runs: the median feature receives about 9.5 effectively distinct labels, and automated judgments find substantial differences in meaning. Variation in held-out AUC does not track label instability, and label indices from different runs share less than a fifth of their top-ten retrieval results. Evidence selection produces the largest contrast in our factorial design, while context diversity correlates with instability (ρ ≈ 0.54). Second, we use these findings to test mitigations. With examples drawn at random, greedy decoding does not reduce median instability, and aggregating runs reduces score SD more than medoid-label instability against matched controls. Selecting the explainer’s top-k examples deterministically roughly halves instability at the same pipeline-run budget and holds on new features, yet even then most features remain above the near-unanimity threshold. These findings motivate treating a feature description as a semantic claim whose uncertainty and dependence on the selected evidence must be characterized alongside its evaluation score.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.