Unlearning During Pretraining Can Improve Tamper Resistance over Data Filtering
Abstract
Preventing the misuse of open-weights language models is a challenging problem, as adversaries have the ability to fine-tune harmful capabilities into models. To make models tamper-resistant with respect to fine-tuning, research has focused on applying unlearning algorithms to a pretrained model or filtering out harmful data from the pretraining corpus. However, post-hoc unlearning methods are mostly not robust to adversarial fine-tuning, and data filtering has steep diminishing returns after the most harmful data is already filtered out. We propose applying common unlearning algorithms during pretraining instead, preventing unwanted capabilities from forming in the first place. Compared to post-hoc unlearning, pretrain-time unlearning significantly improves resistance to adversarial fine-tuning. Pretrain-time unlearning improves over further data filtering in the regime where the most obvious documents have already been removed. Our results highlight that unlearning during pretraining may be an effective tool to prevent undesired capabilities in language models, providing a productive use for harmful data that is usually filtered out.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.