acceptodds
Under review as a conference paper at ICLR 2027

A Harmful Request Is Still a Request: Benign Fine-Tuning Rewrites How Aligned Models Begin to Answer

Abstract

Fine-tuning an aligned language model on harmless data weakens its refusals of harmful requests, and we find this in all 17 open models we test. Two accounts are common, forgetting and shallow alignment, but neither says which harmless data do the damage: two datasets learned equally well can have very different effects on safety. We show that benign fine-tuning, like alignment before it, teaches a model how to begin answering a request, and a harmful request is a request too. The damage therefore sits on the first word of the answer, computed at the last token of the chat template, a position we call the opening, while the model's judgement of harm survives: made to begin with "I", which can open either a refusal or an answer, a fine-tuned model still refuses harmful requests far more often than harmless ones. This tells us where to protect. The opening hold, a loss term that keeps the opening's internal state at the aligned model's value during fine-tuning, cuts harmful compliance from 22.6% to 0.9% over six models and protects across models, datasets and full fine-tuning, whereas freezing the layers thought to carry safety does not; the opening also warns of the damage early, in one forward pass. Benign fine-tuning does not erase what a model knows about harm; it rewrites how the model begins to answer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.