What Can Activation-Patching Evidence Establish?
Abstract
Successful activation patching does not necessarily establish the functional claim it is used to support. Even when expected outputs and measured responses are restored, unmeasured claim-changing alternatives can remain indistinguishable. We formalize this gap as target-relative observability: relative to a declared class of admissible responses, evidence is sufficient only when every response consistent with the observations agrees on the target. On finite-dimensional linear response spaces, we characterize claim-changing blind directions, minimum target-specific measurement completion, and sharp error amplification; an activation-specific corollary shows when a finite probe panel cannot certify unrestricted local response preservation. Prospectively, among center-restoring GPT-2 patches, full-gradient discrepancy predicts withheld finite-response failures better than public-direction screening by 0.080/0.055 within-layer AUC on IOI/GT, with recurring but model-dependent gains across six external model families. Controlled studies show that variance, generic geometric matching, and span coverage can each misrank evidential value, while reference queries can transfer shared nonlinear structure. Reliable validation need not recover the whole response; it must rule out the admissible alternatives that can change the claim.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.