Emergent Misalignment from Safety-Filtered Training Data
Abstract
Recent research has shown that fine-tuning large language models on narrowly misaligned data can lead to broadly misaligned behavior, a phenomenon called emergent misalignment (EM). In this work, we expand these findings by showing that even ordinary training corpora contain examples that can induce EM. In particular, these examples pass common content-level safety filters and therefore appear to be benign. We identify such examples in LMSYS-Chat and Dolci datasets by scoring their representations along directions derived from known EM-inducing and aligned data. Fine-tuning Qwen2.5-7B-Instruct on 1,000 selected safety-filtered examples out of 775k in LMSYS-Chat yields a 15.95% EM rate, compared with 1.70% for a matched random selection. We further validate the effect on 2,190 evaluation questions spanning six behavioral categories, extending beyond the small evaluation suites used in earlier exploratory studies. To characterize the representational signal underlying data selection, we repeat the selection using only 9–12 sparse autoencoder (SAE) features. This recovers 81–90% of the examples selected using the full set of features and induces similar EM rates after fine-tuning. These results show that content-level safety filtering does not guarantee safe fine-tuning outcomes, while model representations provide useful signals for identifying EM-inducing data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.