acceptodds
Under review as a conference paper at ICLR 2027

Self-Projection Distilled Unlearning

Abstract

Reasoning models can generate long chain-of-thought (CoT) traces before a final result, which brings a new risk surface for machine unlearning: hazardous knowledge could be reconstructed step-by-step within the think block even if the final answer is suppressed, and this reasoning can be exposed in agentic traces, audit pipelines, or distilled releases. We find that conventional unlearning methods either fail to forget, or do so at the expense of collapsing benign utility, as the output-level supervision can be hard to affect the unconditional generation mode that drives the CoT remains intact. A key inference-time analysis reveals a different route: contrastive decoding between a safety-prompted and a knowledge-prompted version of the same model can encourage safe behavior without any weight update, indicating that the safe direction is already encoded in the parametric space. We therefore reformulate unlearning as self-consistency resolution: making the unconditional generation mode agree with the model's own safety projection. We instantiate this principle as Self-Projection Distilled Unlearning (SPDU), which distills the model's contrastive logits into its weights through a co-evolving target with the student during training. This converts the inference-time effect into a persistent weight-level change without prompt overhead, and replaces the distribution-shift regret of fixed-teacher distillation with a self-referential bound at the consistent fixed point. Comprehensive experiments validated that SPDU drives unlearning accuracy to near-zero while preserving general capabilities for reasoning tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.