Where Features Are Written and Read: Localizing Native Transformer Channels for Residual Directions
Abstract
Sparse autoencoders, probes, and steering methods find directions in a language model's activations that track concepts, and editing along them can steer the model's behavior. A direction does not say which of the model's own components build the feature or which carry it to behavior, so it is unclear whether an edit along the direction reflects how the model itself computes. We introduce feature-direction localization, which maps a supplied direction to the model's own multilayer perceptron (MLP) channels on both sides of it: writers, the channels that construct the feature, and consumers, the channels through which it reaches behavior. Across interpretable features in several language models, as little as 0.1% of the candidate MLP channels construct 90% of what the MLP layers contribute to a feature, and 0.06% of downstream channels account for 91% of the behavioral effect of steering along it. Connecting the two maps shows that the writers' behavioral effect depends on the same consumers that carry steering: holding those consumers at their clean activations removes much of it. An edit along the direction matched to the writers' change in the feature tests whether steering reproduces their behavioral effect: in one model the matched edit reproduces little of it, while the part of the writers' change that leaves the feature's coordinate unchanged reproduces 85%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.