Sparse-Dictionary Causal Claims Are Not Robust to How They Are Measured
Abstract
Sparse dictionaries (SAEs and crosscoders) are read causally: ablate a feature group, observe the change, compare against a random control. We show that this comparison is under-specified, and measure what each unstated choice does to it. We identify **six failure modes** and size each (Table 1). Dose costs most: supplying a crosscoder's counterpart activation appears to raise an ablation's effect , but at matched norm the encoding is worth and the rest is dose. We prove that whether a variance-matched control can be a population at all is fixed by the activation spectrum before any sampling, and find that in seven of ten layer–side cells of our study none can. On a public GPT-2 SAE, a broadcast direction differs from the group's actual ablation by an order of magnitude, and with the actual ablation the choice between two standard controls flips the sign of the learned-versus-control gap. Across GPT-2 and three Pythia checkpoints the dictionary's *training method* decides whether the comparison resolves (L1 in of cells, TopK in of ). We package the checks as `tabdiff.diagnostics.audit` and measure its limits: against a known nonlinear model, admission alone barely improves a comparison's sign (our registered test of it failed); read afterwards, admission with a resolved interval made no sign error. Our tabular crosscoder is the worked example: its shared group is less sensitive than a matched control, also on six datasets never used for evaluation, at a dose we cannot validate. The usable conclusion is a reporting standard: name the group rule, aggregation, training method, matching covariate and dose, and check the control's diversity ceiling before building it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.