acceptodds
Under review as a conference paper at ICLR 2027

When Unlearned Models Reason, Where Does the Answer Come From

Abstract

Machine unlearning is evaluated by asking an unlearned model about the removed information, but if the model is allowed to reason first, it often answers these forgotten questions correctly again. However, existing evaluations of unlearning only compare accuracy with and without reasoning or inspect reasoning traces for leaked answers, so they cannot show whether the final answer takes its answer from the trace. We instead treat the reasoning trace as an object of intervention, storing and editing it before a separate answering step. Unlike accuracy comparisons and leakage checks, these edits show where the final answer takes its answer from. Concretely, following studies of factual recall, we compare reasoning with a dummy trace, a meaningless text of the same length, and split the recovery measured against it into a dummy penalty and a gain over direct answering. We also mask the answer span against equal-size placebo masks, substitute another author's name of equal token length, and give identical traces to different checkpoints. Experiments on TOFU with seven unlearning methods show that most of the recovery measured against the dummy trace comes from the dummy trace lowering accuracy, not from reasoning. In smaller models, the final answer takes whatever answer the trace states, even in a retrained model that never saw the facts, and unlearning changes what the trace says rather than how the answer uses it. In larger models, unlearning also weakens how much the final answer uses the trace.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.