Bidirectional Causal Tracing and Intervention: Causal Mechanism-Guided Machine Unlearning
Abstract
Large language model unlearning is commonly evaluated through changes in answer probabilities or output behavior, but such behavioral suppression does not demonstrate that the causal support underlying target knowledge recall has been weakened. Reliable unlearning therefore raises two coupled questions: where is this causal support located, and how should it be weakened? To address the first, we use causal tracing under forward and reverse queries to localize internal components whose restorative effects persist across both retrieval directions. To address the second, we optimize localized adapters to directly reduce these components' causal restoration effects, while directional balancing and retention constraints limit collateral damage to non-target knowledge. We call this framework Bidirectional Causal Tracing and Intervention (BCTI). On 50 target facts with Pythia-6.9B-deduped, BCTI achieves an average 87.8% reduction in target answer scores, while a held-out causal audit records a 98.6% ± 1.5% reduction in causal restoration effect. Applying the same procedure to OLMo-7B yields a 97.8% ± 1.9% causal restoration reduction. These results support weakening and auditing causal recall support as a more reliable basis for assessing forgetting than behavioral suppression alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.