Safety Offset Learning: When Benign Fine-Tuning Misaligns LLMs
Abstract
Safety risks from fine-tuning are commonly characterized by the harmfulness of the training data: harmful fine-tuning is considered dangerous, while benign fine-tuning is often assumed to preserve alignment. We argue that this view misses a more fundamental factor—the behavioral transition induced by supervision. Harmful fine-tuning and safety-degrading benign fine-tuning can both be understood as learning the same refusal-to-compliance shift, but on different data distributions. We study whether such a shift learned on benign inputs can generalize to harmful requests. We construct benign examples that an aligned model initially tends to refuse, while retaining harmless target responses that require answering instead. Fine-tuning on these examples increases harmful compliance, particularly when the discrepancy cannot be resolved through narrow surface-level shortcuts. Building on this observation, we introduce SafeOff, a benign rewrite dataset that systematically induces such behavioral shifts. Across multiple aligned LLMs and safety benchmarks, SafeOff substantially weakens safety alignment despite containing no explicit harmful supervision. These results suggest that fine-tuning safety depends not only on what data a model sees, but also on which behavioral boundary the supervision moves.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.