Exposure Leaves a Trace A Provenance Audit for Machine Unlearning
Abstract
We introduce a provenance audit that tests whether an unlearning method removed the of training on the forget examples, rather than only suppressing their outputs. Behavioral parity with the masked reference cannot distinguish suppression from removal. The audit has two stages. First, the exposure-subspace stage constructs a stream-matched training contrast: two models start from one checkpoint and follow the identical training stream, differing only in the loss weight of the forget examples. We define their activation difference as the exposure effect under this matched training intervention, which controls for shared initialization, data order, and training randomness. Removal and insertion show that a low-rank exposure subspace causally affects target behavior under the specified interventions across Llama-3.2-1B, Qwen2.5-1.5B, and Qwen2.5-7B. In the controlled setting, the exposure subspace does not show a pooled advantage over a target-predictive control, as expected when target-correlated activation can coincide with the exposure effect because the forget examples are the target's only source. On pretrained models, where target-predictive directions can also reflect base knowledge, the exposure subspace transfers more behavior. Across seven core unlearning objectives, all six that reach behavioral parity leave a post-unlearning residual whose transplantation into the masked reference recovers 0.15–0.42 of the suppressed behavior. The residual norms remain comparable in scale to retain-only and replace-relabel controls that do not reach parity. An RMU variant supplied with oracle entanglement labels reduces the measured audit diagnostics to the matched-random level, albeit with a large activation displacement. Behavioral parity and removal of the measured training effect are therefore distinct outcomes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.