acceptodds
Under review as a conference paper at ICLR 2027

When Do Explanations Predict Model Changes?

Abstract

Can an explanation predict a model's response to an internal change it has not observed? Matching the original outputs is insufficient: two models can agree on every input yet react differently to the same internal change. We study what finite measurements must retain to predict specified changes, and what to improve when those predictions fail. In a linear response space, a necessary and sufficient condition distinguishes inaccessible information from information discarded by compression; a stability bound separates both from unreliable recovery. Controlled experiments make this distinction actionable: ordinary least squares recovers local responses, whereas principal component analysis (PCA) can lose the required differences despite preserving almost all variance. Retaining those differences repairs prediction. Larger MLPs and pretrained GPT-2 show why local recovery is only part of the problem: new inputs and finite changes introduce different errors. In GPT-2, measuring more inputs improves prediction when local fitting is already accurate. Finite attention-head replacements expose a sharper trade-off: interaction measurements greatly improve predictions of unqueried combinations on the fitting inputs, yet can lose to simpler measurements on more inputs at the same budget. Together, the theory and experiments distinguish when to change the measurements, the representation, or the inputs used to predict model edits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.