acceptodds
Under review as a conference paper at ICLR 2027

Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit

Abstract

One approach in mechanistic interpretability explains model behavior through circuits: the components and connections that carry it. Frozen discovery often returns hundreds of edges, making them hard to inspect, compare, or verify exhaustively. We introduce Circuit Condensation, which post-trains models to concentrate behaviors into smaller causal graphs. Each round prunes low-attribution edges and trains a low-rank adapter to match the original through what remains. It keeps the cut only if task accuracy and held-out perplexity remain within tolerance. Across four behaviors and eight models, condensed circuits are smaller than the strongest published frozen-weight baseline in 30 of 32 settings, by 8.1× on average and up to 316×. Repeating the search without weight updates produces larger circuits in 29 of 32 settings, showing that weight updates drive the reduction rather than search alone. For 19 small circuits, exhaustive subset tests show that 11 are exactly minimal and identify removable edges in the rest. Pair ablations reveal dependencies that single-edge tests miss. On indirect object identification, condensation isolates 24 heads, including 17 with documented roles, whereas a matched frozen circuit needs 61 heads, 36 without documented roles. The result is a sufficient subcircuit of the published mechanism rather than a reconstruction of it. The resulting circuit tracks the original model's next-token distribution and, at matched edge count, predicts its errors better than a frozen circuit in 10 of 11 evaluable settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.