Abstraction of Intermediate Representations for Mechanistic Interpretability via Shapley Indirect Effects
Abstract
Providing interpretable explanations of neural networks is a central challenge in modern machine learning. We study how semantically meaningful changes in input attributes propagate through intermediate representations to affect model outputs, using causal mediation analysis and interchange interventions. Explaining large systems may require abstracting many low-level components into a smaller number of meaningful mechanisms, motivating the abstraction of intermediate representations. We first introduce a Shapley value-based indirect effect that treats causally non-ordered mediators as players in a cooperative game, averages each mediator’s marginal contribution across coalitions, and satisfies the desired additive decomposition property. Using this definition, we formalize abstraction quality in terms of completeness, requiring each input attribute’s effect to be concentrated in a dominant high-level mediator, and distinctness, requiring different input attributes to be associated with different dominant mediators. We then propose a greedy algorithm for constructing the abstraction. Finally, experiments on GPT-2 small and Qwen3.5-9B-base show that clear abstractions emerge at early and intermediate layers, with weaker completeness at later layers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.