Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit
Abstract
One approach in mechanistic interpretability explains model behavior through circuits: the components and connections that carry it. Frozen discovery often returns hundreds of edges, making them hard to inspect, compare, or verify exhaustively. We introduce Circuit Condensation, which post-trains models to concentrate behaviors into smaller causal graphs. Each round prunes low-attribution edges and trains a low-rank adapter to match the original through what remains. It keeps the cut only if task accuracy and held-out perplexity remain within tolerance. Across four behaviors and eight models, condensed circuits are smaller than the strongest published frozen-weight baseline in 30 of 32 settings, by 8.1× on average and up to 316×. Repeating the search without weight updates produces larger circuits in 29 of 32 settings, showing that weight updates drive the reduction rather than search alone. For 19 small circuits, exhaustive subset tests show that 11 are exactly minimal and identify removable edges in the rest. Pair ablations reveal dependencies that single-edge tests miss. On indirect object identification, condensation isolates 24 heads, including 17 with documented roles, whereas a matched frozen circuit needs 61 heads, 36 without documented roles. The result is a sufficient subcircuit of the published mechanism rather than a reconstruction of it. The resulting circuit tracks the original model's next-token distribution and, at matched edge count, predicts its errors better than a frozen circuit in 10 of 11 evaluable settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.