A Reporting Standard for Sparse Autoencoder Ablation Scores
Abstract
Ablation scores for sparse autoencoder latents disagree across dictionaries partly because each dictionary chooses the token at which its own latent is scored. A latent is scored by switching it off at one token and measuring how much the model's output changes, and by convention that token is where the dictionary's latent fires hardest. We isolate the effect with autoencoders trained from a shared initialisation, each differing in a single fitting choice. Across seven base models from four families, the choice of token carries at least 46.6% of the variance in the measured effect, about as much as the choice of latent. We propose a reporting standard: record the measurement position, and when dictionaries are compared, fix it outside all of them. Fixing the tokens makes dictionaries rank latents more consistently, with a bootstrap interval excluding zero on every model at the main corpus size. Both items take a line of evaluation code, and the second needs no more forward passes than letting each dictionary choose. Left to choose, dictionaries rarely agree on where to measure a latent, and agree less as the evaluation corpus grows. Released Qwen-Scope and Gemma Scope dictionaries measure most latents matched across a pair of dictionaries at different tokens, including many that the pair encodes almost identically. Code: https://anonymous.4open.science/r/sae-artifact-8E6D/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.