acceptodds
Under review as a conference paper at ICLR 2027

Understanding Safety Unalignment through Multi-Stage Reinforcement Learning Attacks

Abstract

Open-weight large language models (LLMs) remain vulnerable to post-training that removes safety alignment. Existing defenses, however, are evaluated predominantly against supervised fine-tuning (SFT), with only limited investigation of their robustness to reinforcement learning (RL) attacks. We first analyze safety unalignment under Harmful GRPO and find that harmful behavior emerges through distinct phases: from explicit refusal, through intermediate responses that repeat or condemn harmful requests, to affirmative prefixes followed by harmful content. We further show that harmful RL substantially degrades general generative utility, a failure mode largely obscured by conventional closed-ended benchmarks but exposed by open-ended evaluations. Guided by these observations, we introduce a simple two-stage post-training attack. First Few Tokens Distillation (FFTD) distills the distribution over the initial response tokens from a weakly unaligned teacher to accelerate safety unalignment, while On-Policy Distillation on a Benign Dataset (OPDBD) subsequently restores utility using only benign prompts and the original model as the teacher. Across multiple open-weight models, our approach achieves high harmful compliance while preserving substantially more utility than conventional harmful fine-tuning and circumvents representative defenses including TokenBuncher and SEAM. Our analysis reveals that these defenses rely on assumptions about rollout distributions or local optimization dynamics that can break under our proposed RL attack strategy, highlighting the need for mechanism-aware evaluation against heterogeneous post-training attacks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.