Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models
Abstract
Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a _primary_ component can inhibit the activation of a _backup_, leading to issues with ranking components. _Actual causality_ studies the structure of such interactions via _witnesses_: variables that provide contextual information to resolve interaction terms. However, estimation with witnesses typically requires combinatorial enumeration and is infeasible in practice. We introduce the _witness-integrated set effect (WISE)_, a family of causal estimands that build on the witness mechanism while taking expectations over _sets_ of causes and witnesses to remain computationally feasible. Building on this approach, we introduce _JuntaLearner_, a gradient-based circuit discovery method that learns to rank components by their causal impact across varying-sized sets of components and witnesses. Alongside faithfulness metrics, we introduce measures of necessity and task specificity, and the _circuit recognition score (CRS)_ to summarize each metric across circuit sizes while emphasizing effects achieved by small circuits. Across tasks and models of increasing size, _JuntaLearner_ achieves higher mean CRS compared to attribution baselines on all metrics. Since its cost does not grow with the number of candidate components, _JuntaLearner_ scales to large models while accounting for set-level interactions and avoiding first-order approximations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.