When Do Neural States Realize the Same Causal Variable?
Abstract
Mechanistic explanations often describe neural states in terms of high-level variables. For such descriptions to function as causal abstractions, the distinctions they omit must be irrelevant to the downstream counterfactual behavior of interest. Yet causal adequacy is usually assessed at a particular neural state, leaving open whether the same description remains adequate as computation unfolds. We study this question in a controlled retrieval task where direct and relational queries reach the same selected entity through different computations. The relevant semantic intermediate is fixed by construction, while we independently vary how the entity is selected, which value it is assigned, and where its complete assignment record appears. We find three stages with different causal organization: entity identity is sufficient at an early direct-query state under our interventions, later selection states also depend on assignment-record arrangement, and later states are organized by answer identity. State-restoration interventions connect these stages causally: an effect established at selection is later reversible through a downstream context-sensitive state. Meanwhile, states reached through the direct and relational procedures become increasingly interchangeable even as the distinctions required for causal sufficiency become finer. Thus causal granularity changes across computation even though the semantic intermediate is held fixed: the same high-level variable can be causally adequate at one stage and too coarse at another. Mechanistic explanations must establish not only which high-level variable a neural state corresponds to, but also over which part of the computation that variable remains causally adequate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.