Feature-Resolved Attention
Abstract
Dictionary learning methods such as sparse autoencoders aim to provide an interpretable, mono-semantic basis for a model's computation. Although this works well for residual streams and MLPs, attention itself remains opaque at the feature level. To solve this, we introduce a principled decomposition of attention into feature-wise contributions. We call the resulting object _Feature-Resolved Attention_ (FRA). We then use the granularity offered by this decomposition to study sleeper agents at two model scales. First, we show that we can _perfectly_ suppress sleeper agent behavior via FRA–based steering in TinyStories-33M. Strikingly, in around 40% of cases we recover the original text word-for-word, well above conventional SAE steering at the same hookpoint. Second, we consider a sleeper agent in an 8B-parameter Llama model. FRA based steering outperforms both conventional steering and Difference-of-Means steering in recovering benign from sleeper sentences. Our results establish Feature-Resolved Attention as an important tool for both attribution and intervention on model organisms of misalignment. Code is available at https://anonymous.4open.science/r/fra-anon-iclr-F346/README.md.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.