acceptodds
Under review as a conference paper at ICLR 2027

Attribution Vergence: A Mechanistic Indicator across Grokking and Adversarial Transfer

Abstract

Connecting internal computational mechanisms to macroscopic model behaviors, such as grokking and adversarial transfer, is one of the central goals of mechanistic interpretation. However, while existing methods analyze isolated components or global representations, whether the structure of attributed contributions supporting model decisions admits a general macroscopic observable remains fundamentally open. We introduce Attribution Vergence, a measure that quantifies this structural reorganization across feature aggregation by measuring the change in R\'enyi-2 effective dimensionality. Under a unified Aggregation Node formulation encompassing attention and convolution, we ground the measurement via attribution preservation, axiomatic attribution compatibility conditions, an information-theoretic interpretation, and analysis on invariance, bounds, stability, and computation. Experiments on grokking and generative adversarial transfer show stronger behavioral responses than prior analysis tools and additional gains from the formulation. The profiles reveal perturbation-dependent response locations, source indicators shared across transfer targets, and attribution reorganization across layers after training accuracy saturates. Together, these results establish a computable measurement framework linking attribution organization to macroscopic behaviors, as a step toward a scientific understanding of AI models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.