acceptodds
Under review as a conference paper at ICLR 2027

: Testing Head Effects Through Averaged Activations via Inject-and-Block

Abstract

Mechanistic interpretability seeks to explain how internal components produce model behavior. After locating a task-relevantant attention head, a central question is whether its effect depends on a shared activation direction or on input-dependent structure. This distinction determines how an averaged direction should be interpreted and reused. We introduce THETA, a framework combining mean-direction dependence testing, write geometry, and donor reuse analysis. Its core protocol, Inject-and-Block, patches a head, measures the induced residual-stream change, and removes its mean-direction projection, with full-write and random-direction controls. Our contributions are threefold: (i) a controlled protocol connecting an individual head's measured effect to its average write direction; (ii) evidence across five configurations spanning Gemma, GPT-2, Qwen, and Llama that heads exhibit distinct mean-direction removal responses, with write alignment distinguishing the resulting classes; and (iii) a Gemma case study identifying query matching as a condition for reusing writes whose effects persist after mean-direction removal. Same-query donors achieve normalized recovery of , compared with for different-query donors, supported by stronger within-query alignment. THETA connects component-level causal tests with the choice between shared directions and query-matched donor writes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.