LLMs Can Unlearn Refusal with Only 1,000 Benign Samples
Abstract
Fine-tuning on harmful data is known to undermine the safety alignment of Large Language Models (LLMs). Similar effects have also been observed when fine-tuning exclusively on benign data. However, this phenomenon has received limited follow-up attention, largely because benign fine-tuning typically induces only modest safety degradation. We revisit benign fine-tuning from a novel refusal unlearning perspective. Our study is motivated by a key observation overlooked in prior work: aligned LLMs predominantly respond to unsafe queries with refusals that often begin with a small, fixed set of prefixes (e.g., “I'm sorry”). We demonstrate that this rigid refusal pattern constitutes a vulnerability and exploit it to fine-tune LLMs in a way that disrupts refusal generation behavior. Surprisingly, we find that merely 1,000 benign samples are sufficient to achieve this goal. We apply this approach to a total of 16 LLMs, including various open-source models from Llama, Qwen, and Gemma families, as well as closed-source models such as Gemini and GPT. Experimental results show consistent and substantial degradation in the safety scores of previously aligned LLMs, significantly exceeding that observed under the benign fine-tuning baseline. Importantly, this vulnerability manifests not only during the supervised fine-tuning stage but also throughout the reinforcement learning phase. Our method is further substantiated by theoretical analyses that establish a connection between the refusal strength of prefixes and the effectiveness of refusal unlearning. Taken together, these findings suggest that current safety alignment may rely heavily on token sequence memorization rather than reasoning, motivating future work beyond simple refusal mechanisms. The code has been released anonymously at: https://anonymous.4open.science/r/refusal-unlearning-2CD7/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.