acceptodds
Under review as a conference paper at ICLR 2027

Capability Changes Under Safety-Alignment Removal: A Stage-Resolved Study of LLMs

Abstract

Harmful fine-tuning can remove the safety alignment of open-weight language models. Studies of such attacks usually report two final numbers: attack success on harmful prompts, and one average utility score compared with the aligned model. We show that this pair can mislead. We follow four models from three families through each stage of a removal pipeline, including branches that start from adapted defenses. At every stage, we measure harmful compliance and task utility separately, and we add matched evaluation controls. Three lessons emerge. First, similar attack success comes with different utility outcomes. Final HarmBench attack success reaches 88–97% for every model. Yet average utility ranges from a small gain to a loss of 9.3 points over three training seeds. Different skills decline at different stages. Second, the reference point decides the verdict. A model can recover strongly from a defended start and still remain below the original aligned model. Third, closing a gap is not the same as explaining it. When we change only how answers are requested, one large math gap disappears, but only because the aligned model's score falls to the trained model's level. The trained model now answers without working, and the new request makes the aligned model do the same. Instruction-following losses persist under every control we test. Claims that a removal preserves capability therefore need comparisons that match task, reference, and evaluation protocol.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.