acceptodds
Under review as a conference paper at ICLR 2027

Separating representational redundancy from causal responsibility

Abstract

Neural network interpretations often treat decodability as evidence of functional use. If a linear classifier predicts a task variable from a component's hidden state, that component may be taken both to contain and contribute the information. This work separates those properties. A component may encode task information the model does not use, while another may control the output without linearly encoding the answer. Two architectural mechanisms organise the distinction. Transport spreads task information between components and determines where it can be recovered. Selection governs which available components affect the output. In two transport sweeps, selection has its largest effect after transport creates usable alternatives but before the original component loses measured responsibility. A follow-up statistic and decision rule for this non-monotonic pattern were fixed after the initial result and confirmed on fresh seeds. Across recurrent, graph, convolutional, transformer, and mixture-of-experts networks, greater transport capacity is accompanied by broader recoverability with little change in task accuracy. Fixed-weight route cuts isolate transport where such interventions are available. In a graph network with its body frozen, the task variable is decoded from components other than its source with 0.88 accuracy, yet source removal reduces the trained readout to chance. Retraining only the readout recovers 0.87 accuracy, showing that the information was available to a new readout but unused by the original one. Conversely, transplanting a pointer node redirects the answer even though the answer remains linearly undecodable there. Representational availability and causal responsibility are distinct and should be measured separately when attributing functions to network components.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.