Transferable LLM Misalignment Attack via Safety Task Vector Extraction
Abstract
Safety alignment is critical to ensure that large language models appropriately respond to harmful instructions while remaining helpful on benign queries. Existing works on misalignment attacks have demonstrated that safety alignment can be easily compromised for open-weight models. However, these attacks are typically constructed for a specific target model and need to be reconstructed when the target changes. In this paper, we introduce a task arithmetic-based attack that extracts a transferable safety parameter direction that can be reused across multiple post-trained models derived from the same pretrained base model. To construct this direction, we first train paired models with and without safety-related data and extract the safety-related task vector through task arithmetic. Then, we utilize multi-seed sign consensus to suppress training noise and retain stable safety-related updates. The resulting safety task vector can be subtracted from various models derived from the same base to weaken their safety behavior. Extensive experiments across multiple model families demonstrate that our method effectively compromises safety alignment while largely preserving general model utility and exhibits transferability among models sharing the same pretrained base.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.