acceptodds
Under review as a conference paper at ICLR 2027

Finding Subspace Circuits with Subspace Attribution Patching

Abstract

Edge attribution patching (EAP) has been widely adopted as a tool for finding circuits — small, task-specific computational subgraphs — within language models. A growing body of evidence suggests that the model components within a circuit communicate via low-rank subspaces of the residual stream. However, while circuit localization is well-studied, the study of the subspaces used by a circuit has received little attention. This work proposes subspace attribution patching (SAP), a finer-grained formulation of EAP, which finds circuit edges jointly with the causally-relevant subspace for each edge. We rigorously evaluate a range of methods for constructing subspace bases for SAP, relying on spectral decompositions of model weights and activations. Furthermore, we demonstrate that several plausible SAP approaches can perform well while not being faithful to the model's computations, falling prey to the Interpretability Illusion of Makelov et al. (2024). We conclude by discussing a method to diagnose this illusion. Our findings enable researchers 1) to jointly localize circuit edges and low-rank information subspaces for a given task, and 2) to detect illusory results, paving the way for more reliable research in this area.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.