acceptodds
Under review as a conference paper at ICLR 2027

Can Circuit Discovery Be Learned? A Graph Neural Network Approach

Abstract

Circuit discovery aims to find the sparse subnetworks responsible for specific behaviors in language models. Existing automated methods either approximate the effect of interventions to first order, as in attribution patching, or optimize a new set of mask parameters for every task, as in Edge Pruning. In both cases, nothing reusable is learned about how to find circuits. We take a different approach and train a graph neural network to predict circuits directly from activations, extending recent work that trains models to interpret model internals. The network operates on the transformer's computation graph, takes clean and counterfactual activations as input, and assigns a score to each edge. Keeping the highest-scoring edges gives a circuit at any target sparsity in a single forward pass, without attribution-derived features or test-time optimization. Scorers trained on individual tasks outperform the previous state-of-the-art methods EAP and EAP-IG on integrated circuit performance (CPR) on the MIB Benchmark. We then train a scorer jointly across multiple tasks and this scorer successfully infers the task and its circuit from the activations alone, remaining competitive against task specific GNN. We see this as the first step toward general-purpose learned circuit discovery.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.