What Goes Up Should Come Down: A Training-Distribution Control for Interpretability Metrics
Abstract
Interpretability metrics are often supported by showing that their results are reproducible across seeds, model conditions, and analytical choices. Such robustness does not establish that a metric is valid for the interpretation attached to it. We demonstrate this distinction using a relation-referenced metric for sparse autoencoder (SAE) features in protein language models. The metric, , measures whether feature activations cluster across residues in spatial contact. In a controlled comparison of masked and causal models, gives a plausible masked-model advantage at every sampled depth. The difference reproduces across model seeds, persists under prediction matching, and survives six diagnostic analyses spanning dictionary fitting, reconstruction quality, residue selectivity, and representation basis. However, after retraining the same models on sequences with native residue order destroyed, increases in 52 of 54 objective-by-depth-by-seed cells. Random-initialisation and no-model baselines reveal complementary problems: the trained masked-model score does not exceed the random-initialisation range, and a cysteine indicator exceeds every learned feature in 32 of the same 54 cells. These results do not rule out structural information in , but they show that its robust masked–causal difference is not specific evidence of structural organisation learned from native residue order. Applying the same framework to further published interpretability metrics reveals different patterns of failure across the training-distribution, random-initialisation, and no-model controls. Interpretability metrics therefore require controls matched to each link between a score and its intended interpretation; reproducibility within a measurement pipeline is not sufficient evidence of construct validity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.