Intent Heads Build the Path to Emergent Misalignment Before Fine-Tuning
Abstract
Fine-tuning large language models on narrowly scoped data can induce broad harmful behaviors beyond the training domain, a concerning phenomenon known as emergent misalignment (EM). While prior studies have linked EM to global malicious personas, what intrinsically causes narrow fine-tuning to induce broadly generalized misalignment remains unclear, limiting our understanding of EM's underlying mechanisms and our ability to prevent it proactively. In this work, we find that pre-existing attention heads responsible for intent understanding, termed intent heads, causally contribute to the formation of EM. By strategically ablating and restoring these heads during fine-tuning, we demonstrate that these heads shape fine-tuning dynamics, biasing optimization toward broadly misaligned outcomes. Furthermore, early ablation of these heads leaves a persistent effect on the subsequent optimization trajectory and fundamentally attenuates misaligned persona features within the model's internal states. Together, these findings highlight an unexplored form of component reuse that occurs during the training stage. Based on these findings, we introduce Restored Intent Head Ablation (RIHA), a lightweight mitigation strategy that temporarily ablates intent heads during early fine-tuning. We show it effectively suppresses EM and can further complement and amplify existing mitigation methods, with zero training overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.