acceptodds
Under review as a conference paper at ICLR 2027

What Can Activation-Patching Evidence Establish?

Abstract

Successful activation patching does not necessarily establish the functional claim it is used to support. Even when expected outputs and measured responses are restored, unmeasured claim-changing alternatives can remain indistinguishable. We formalize this gap as target-relative observability: relative to a declared class of admissible responses, evidence is sufficient only when every response consistent with the observations agrees on the target. On finite-dimensional linear response spaces, we characterize claim-changing blind directions, minimum target-specific measurement completion, and sharp error amplification; an activation-specific corollary shows when a finite probe panel cannot certify unrestricted local response preservation. Prospectively, among center-restoring GPT-2 patches, full-gradient discrepancy predicts withheld finite-response failures better than public-direction screening by 0.080/0.055 within-layer AUC on IOI/GT, with recurring but model-dependent gains across six external model families. Controlled studies show that variance, generic geometric matching, and span coverage can each misrank evidential value, while reference queries can transfer shared nonlinear structure. Reliable validation need not recover the whole response; it must rule out the admissible alternatives that can change the claim.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.