Implicit Evocation in Unlearning-Based Alignment: Exploration and Explanation
Abstract
Language-model unlearning is commonly evaluated at the output, leaving open whether reduced target answering is accompanied by representational erasure. We study this question on RWKU using two instruction-tuned, approximately 4B-parameter models, three unlearning objectives (GA, NPO, and ATU), and both full-parameter and LoRA updates. Applying the model’s shared unembedding as a raw logit lens, we find that target-answer subtokens can remain more highly ranked than task-irrelevant baseline subtokens at intermediate layers even when their final-layer ranks decrease. We call this tested pattern implicit evocation. The observation establishes residual linear decodability under an oracle target-subtoken readout; it does not by itself establish attack-level recoverability or universality across unlearning methods. Under an idealized residual-block model and explicit boundedness and contractivity assumptions, we also derive conditions that favor stronger late-layer updates. This conditional analysis is consistent with late-layer suppression but does not fully explain the observed non-monotone mid-layer peak.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.