Capability Changes Under Safety-Alignment Removal: A Stage-Resolved Study of LLMs
Abstract
Harmful fine-tuning can remove the safety alignment of open-weight language models. Studies of such attacks usually report two final numbers: attack success on harmful prompts, and one average utility score compared with the aligned model. We show that this pair can mislead. We follow four models from three families through each stage of a removal pipeline, including branches that start from adapted defenses. At every stage, we measure harmful compliance and task utility separately, and we add matched evaluation controls. Three lessons emerge. First, similar attack success comes with different utility outcomes. Final HarmBench attack success reaches 88–97% for every model. Yet average utility ranges from a small gain to a loss of 9.3 points over three training seeds. Different skills decline at different stages. Second, the reference point decides the verdict. A model can recover strongly from a defended start and still remain below the original aligned model. Third, closing a gap is not the same as explaining it. When we change only how answers are requested, one large math gap disappears, but only because the aligned model's score falls to the trained model's level. The trained model now answers without working, and the new request makes the aligned model do the same. Instruction-following losses persist under every control we test. Claims that a removal preserves capability therefore need comparisons that match task, reference, and evaluation protocol.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.