Activation Steering Induces Emergent Misalignment: Steering Vector Construction and Comprehensive Evaluation
Abstract
Activation steering (AS) has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs). AS modulates model behavior in a flexible and lightweight manner by injecting task-specific steering vectors (SVs) into intermediate activations. Meanwhile, recent work has identified emergent misalignment (EM) as an emerging safety concern, whereby finetuning on unsafe examples from a narrow task unexpectedly induces broadly unsafe behavior in unrelated domains. With the increasing popularity of AS, it is pressing yet under-explored to thoroughly investigate whether AS can induce EM. In this paper, we conduct a comprehensive study of activation-steering-induced emergent misalignment. Specifically, (1) we identify a steering vector construction to confirm that SVs derived from a narrow task can indeed induce misalignment across broad unrelated domains. Qualitative measurement further reveals that AS-induced EM may pose more severe safety risks than finetuning-induced EM, i.e., AS elicits more actionable harmful responses with greater semantic coherence than those generated by finetuned models. (2) We characterize the properties of AS-induced EM by analyzing key steering-specific factors, including steering magnitude, the low-rank structure of the steering subspace, and the number of fine-tuning epochs used during steering-vector construction. (3) We evaluate the robustness and sensitivity of AS-induced EM across diverse model families, model scales, target tasks, and intervention layers. Our findings reveal activation steering as a significant yet underexamined source of emergent misalignment and provide an activation-space perspective on emergent misalignment and its safety risks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.