Weight-Dictionary Decomposition: Reading a Transformer's Residual Stream with Its Own Writes
Abstract
Sparse autoencoders interpret the residual states of a transformer through an overcomplete dictionary trained on activations. We build the dictionary from the weights instead. Each component's write into the residual stream is confined to a direction or subspace that its weights fix in advance, so collecting those directions from the checkpoint yields an overcomplete, training-free dictionary in which each atom is labeled with the component that wrote it. Weight-dictionary decomposition (WDD) sparse-codes an observed state over this dictionary to infer its largest native writes and their strengths. We score it against exact ground truth from hooked forward passes in eight models from six families, up to 7B parameters, on two corpora. An identified write has the right sign in more than 99.7% of cases in all eight models. Its magnitude is recovered with a median relative error between 0.25 and 0.54 in seven of them, biased upward for small writes because a 64-atom support absorbs the writes it leaves out, and it is not recovered in Pythia-6.9B. The identification rate rises with scale to 0.98 for Llama-7B. The rest of the support is not a list of writers, since only a quarter to seven tenths of the MLP atoms it selects are real writes at that token. Identification is bounded by an observability ceiling that we measure in all eight models. Once the other writes cancel a write along its own direction, OMP no longer finds it at the tested budget, although a probe on the full state shows that the information usually survives the cancellation. Used as measurement axes, the write directions single out a GPT-2 neuron that appears to erase the model's massive activation channel, and interventions under a frozen-normalization control reinterpret it as a participant in maintaining that channel, one that a distributed suppression field keeps below its clean level. In pre-specified comparisons at thirty nominees per arm, WDD beats a marginal outlier screen at nominating live circuits but does not outperform a correlation screen, and we argue that its distinctive contribution is the mechanistic content that each nomination carries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.