Self-Discovering Circuit Models
Abstract
*Circuit discovery* is a central task in mechanistic interpretability, aiming to identify concise sets of transformer components responsible for a given behavior using either task-agnostic patching (e.g., zero, mean, resampling) or task-specific counterfactual patching. Obtaining concise circuits via *post-hoc* discovery methods is computationally demanding, while faster approximations often produce suboptimal, less faithful circuits. Moreover, patching combines activations in ways the model never encounters during training, which can distort the recovered circuits. To mitigate these limitations, we propose a self-supervised fine-tuning approach yielding *Self-Discovering Circuit Models* (SDCMs): transformers trained under task-agnostic patching to produce input-specific circuits for their own predictions as an integral part of their output, while preserving comparable predictive performance. Aggregated across examples, these input-level circuits can form a single circuit for an entire behavior. Empirically, under task-agnostic patching, SDCMs discover circuits far more efficiently than post-hoc methods, and these circuits are substantially more faithful and concise, with a single circuit outperforming post-hoc circuits extracted separately for each patching scheme. Moreover, although counterfactual patching is undefined on general training text, our training reshapes the model parameters so that post-hoc methods recover more faithful and concise circuits from SDCMs than from conventionally trained models under such patching. Lastly, we find that SDCM circuits remain superior to post-hoc circuits even on a conventionally trained model that never received SDCM training, indicating that their advantage lies also in the circuits themselves, and not only in the trained model. This opens the door to using SDCMs as “twin models” that efficiently supply concise and faithful circuits for their conventionally trained counterparts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.