Demystifying Variance in Circuit Discovery of Large Language Models
Abstract
Circuit discovery is a key technique in mechanistic interpretability to pinpoint the model components that are crucial for performing a given task. Although the current state-of-the-art method (EAP-IG) performs well on the metric of (un)faithfulness, it suffers from substantial variability. This includes _resampling variance_, where the circuit changes when we probe with a new batch of data from the same distribution; _phrasing variance_, where the discovered circuit shifts when the prompts are rephrased; and _sample-wise variance_, where a circuit with low population unfaithfulness exhibits large fluctuations in unfaithfulness across individual samples. This paper studies the roots of these variances. We demonstrate that conductance-based EAP (CEAP), a more mathematically principled alternative to EAP-IG, lessens resampling variance. We further show that phrasing variance arises because prompts expressed in different templates tend to activate different circuits with limited cross-template generalizability, and that weight sparsity can mitigate this limitation. Regarding sample-wise variance, we find that extremely poor unfaithfulness scores often arise from how the metric is normalized, rather than from large discrepancies between the circuit and the full model's behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.