Sparse-Autoencoder Attribution Graphs Are Decided at the Margin
Abstract
Attribution graphs over sparse-autoencoder features predict an ablation’s effect, and are compared across seeds. Under a TopK rule a feature is active where its pre-activation is positive and among the largest, and its signed margin, positive on an active feature, measures the move that changes the membership. On five public TopK pairs of three Pythia models, zeroing one feature moves a downstream feature across the cutoff in 0.74 of cases or more, where the crossing features carry most of the response’s energy. A write scales a feature’s code by its dose. Over three doses the graph held on the active set has a median relative error of 0.247 or more on the code change. The margin graph adds the write’s first-order change to every pre-activation and re-ranks, at the same one Jacobian-vector product per live feature, and errs by at most 0.033. At the median a crossing coordinate’s code moves several times as far as its pre-activation, and the crossing features’ decoded change has 1.24 times the model’s change in norm or more. On one testbed of eight runs, sparse nonnegative maps writing one dictionary’s atoms over another’s explain 0.85 of the energy of the operator between two layers’ codes on the source’s block of live upstream and downstream features, and 0.47 on graphs masked by their margins’ signs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.