acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Circuit Discovery with Learnable Corruptions

Abstract

Patching-based circuit discovery has a hidden degree of freedom: the corruption that defines the behavioral contrast. Existing methods fix this corruption and optimize only the sparse circuit, even though a patching score is not a property of a circuit alone, but of a circuit under a chosen behavioral contrast. We make this dependence explicit by recasting circuit discovery as a joint optimization over sparse circuits and reference-anchored admissible corruptions. In our reformulated problem, the benchmark reference is retained as an anchor for constructing semantically admissible and behaviorally disruptive counterfactual families, with fixed-corruption discovery recovered as a boundary case. We instantiate this formulation with SEED (**S**tructure **E**dit and **E**dge **D**iscovery), a tractable alternating procedure that updates the corruption coordinate under the current circuit and updates a shared sparse circuit under the resulting contrasts. Across eight circuit-discovery tasks, under both denoising (sufficiency) and noising (necessity) discovery objectives and with both Faithfulness and Completeness measured on held-out data, learning corruptions improves over optimizing the circuit alone, although the gains are task-dependent. Yet the corruptions selected by SEED do not transfer uniformly to non-SEED methods, indicating that they behave as circuit-coupled best responses rather than universally better corruptions. These results suggest that circuit discovery should be understood as a coupled problem over circuits and behavioral contrasts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.