acceptodds
Under review as a conference paper at ICLR 2027

Demystifying Variance in Circuit Discovery of Large Language Models

Abstract

Circuit discovery is a key technique in mechanistic interpretability to pinpoint the model components that are crucial for performing a given task. Although the current state-of-the-art method (EAP-IG) performs well on the metric of (un)faithfulness, it suffers from substantial variability. This includes _resampling variance_, where the circuit changes when we probe with a new batch of data from the same distribution; _phrasing variance_, where the discovered circuit shifts when the prompts are rephrased; and _sample-wise variance_, where a circuit with low population unfaithfulness exhibits large fluctuations in unfaithfulness across individual samples. This paper studies the roots of these variances. We demonstrate that conductance-based EAP (CEAP), a more mathematically principled alternative to EAP-IG, lessens resampling variance. We further show that phrasing variance arises because prompts expressed in different templates tend to activate different circuits with limited cross-template generalizability, and that weight sparsity can mitigate this limitation. Regarding sample-wise variance, we find that extremely poor unfaithfulness scores often arise from how the metric is normalized, rather than from large discrepancies between the circuit and the full model's behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.