The Role of Learning and Memorization in Relabeling-Based Unlearning
Abstract
This work studies how the nature of a response generated by a model impacts the efficiency of relabeling-based unlearning, a common unlearning technique that trains the model to fit an unlearn set (i.e., a dataset that we wish the model to unlearn) with alternative responses to prevent it from generating unwanted outputs that align with the unlearn set. We distinguish between two different ways models can generate undesirable outputs: learning-based generation, where the model learns an underlying rule connecting the input and the response (e.g., social stereotypes), and memorization-based generation, where the model memorizes specific information about a given input (e.g., private information like a phone number). We demonstrate through binary classification and TOFU-LM experiments that relabeling-based unlearning converges more slowly and causes greater temporary retain-accuracy loss for learning-based generation than for memorization-based generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.