How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
Abstract
The circuits framework in mechanistic interpretability aims to identify sparse subgraphs of model components that are causally responsible for a behavior, typically evaluated by measuring _necessity_ and _sufficiency_. But these criteria say little about whether a circuit consistently captures how a model performs a task, or if it is specific to that task. We study these two properties, _consistency_ and _specificity_, across six tasks and five models, extracting circuits at the component level (attention heads and MLP blocks) and at the level of individual MLP neurons. We find that component-level circuits are highly consistent and causally important on most tasks, but they are not specific: ablating one task's circuit damages another task's performance about as much as that task's own circuit does. Neuron-level circuits, on the other hand, exhibit higher task-specificity but are far less consistent within tasks. This is explained by circuit overlap: component-level circuits share most of their components across all task pairs, related or not, while neuron-level circuits overlap only between closely related tasks. In a case study of the components shared by the task circuits of `Llama-3.2-3B` we show that they consist mostly of MLP blocks, we show that they consist mostly of MLP blocks, while the few attention heads within turn out to have general-purpose roles. Overall, our findings raise questions about the degree to which circuits can support targeted understanding of, and intervention on, model behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.