Behavioral forgetting does not necessarily erase causal accessibility
Abstract
When a network’s output on an input changes over training, it is natural to describe the network as having “forgotten” whatever computation previously produced the old output. This description conflates two distinct properties. We demonstrate that behavioral state, causal accessibility, and behavioral recoverability are empirically dissociable: a network can produce behavior fully consistent with having overwritten an association while retaining, in an intervention-accessible form, the causal capacity to reinstate the original output. This is a claim about what the interventions establish – that a historical direction in hidden-state space remains causally load-bearing – not a claim that a specific “old computation” or “circuit” has survived. Controls rule out generic perturbation sensitivity, lineage-independence, diffuseness, construction-arbitrariness, and pure fc2-geometry explanations. Upstream analyses (fc1 weight-space substrate; natural indirect effect decomposition) confirm the persistence is genuine upstream-of-readout structure. The dissociation holds across 10 model configurations including Transformers. Regarding mechanism: the evidence is more consistent with partially overlapping representational resources than with two cleanly independent pathways – but this interpretation is not uniquely established by our interventions, and we make no claim about the identity of the surviving internal structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.