Let's interpret step-by-step
Abstract
Interpretability research often seeks to explain model computation using ordered sequences of discrete steps, particularly when describing latent reasoning pipelines or circuits. When selecting features for circuits, we should therefore prefer features that actually capture discrete steps in the model's forward pass. In practice, however, this criterion is rarely verified or explicitly targeted when building feature dictionaries. We propose a new metric, *punctuality*, to measure if a feature was computed in a step, rather than a gradual refinement process. As a case study, we analyze a popular sparse autoencoder (SAE) feature circuit and find that many features are punctual and represent a shared step in computation. Next, we assess residual stream SAE features and find them to be less punctual than neuron basis features—and furthermore, that more punctual SAE features are less causally important, limiting their value for interpreting latent reasoning mechanisms. Finally, we show that crosscoders, which are designed to merge features that persist across layers, outperform SAEs in finding features that are both punctual and causally important. By explicitly measuring punctuality along with other desiderata like causal importance and interpretability, we can assess which features can be used in circuits or explanations of latent reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.