acceptodds
Under review as a conference paper at ICLR 2027

Aggregation Order in Activation Patching: Which Circuit Is Selected and What It Preserves

Abstract

Activation-patching evidence supports a mechanistic claim only as far as the measurement pipeline behind it can distinguish that claim from alternatives. Here we study one step of that pipeline: how effects across prompts are aggregated. A circuit is often called faithful when ablating everything else preserves a task score aggregated over prompts, but opposite-signed prompt effects can be combined before or after taking their magnitude. We isolate this aggregation order: within each comparison, the model, candidate universe, prompt set, circuit size and exhaustive search are held fixed, and we measure which circuit each order selects and what that circuit preserves on held-out prompts. Across twelve independently drawn discovery-and-evaluation runs per model from one template bank, in the published indirect-object identification (IOI) universes of GPT-2 small and Pythia-160m, under mean ablation, averaging first selects circuits with 1.73× and 3.65× the full-ablation-normalized held-out KL of those selected by per-prompt-absolute scoring (raw gaps 0.049 and 0.129 nats; the two models' ablation scopes differ). The gap is already present on the discovery prompts. On held-out prompts, most of the excess lies in answer-set probability mass and, on normalized KL, average-first selections were not established to differ from random subsets of the same size. In Pythia, average-first circuits keep the clean model's top token on 75.6% of prompts versus 89.7% for per-prompt-absolute circuits (a prespecified descriptive endpoint), despite similar binary task accuracy. Under counterfactual patching the ratios fall to 1.11× and 1.12×, and the registered operator interaction was confirmed on both models. A preregistered six-family study confirms the mean-ablation contrast, on the relative scale, in only two GPT-2 families, and a preregistered 8B layer-group extension missed its materiality gate. We propose neither a new discovery algorithm nor a universally preferable faithfulness score: changing only the aggregation order can change which circuit is selected and what it preserves, so faithfulness claims should state their aggregation and intervention protocol and evaluate directly the property they are intended to preserve.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.