acceptodds
Under review as a conference paper at ICLR 2027

GraSafe: Mitigating Gradient Conflicts for Preserving LLM Safety during Fine-tuning

Abstract

Aiming to adapt Large Language Models (LLMs) to diverse downstream tasks, fine-tuning is typically performed on task-specific datasets. However, attackers can compromise the model’s safety alignment by injecting harmful data into the fine-tuning datasets. More concerningly, recent studies have shown that safety degradation can also occur even when the model is fine-tuned exclusively on benign data. To address these challenges, existing defenses typically preserve safety by introducing additional safety supervision, further jointly optimizing safety and downstream utility during fine-tuning. Nevertheless, such strategies fail to explicitly distinguish utility-relevant updates from safety-conflicting ones, potentially limiting their ability to achieve a favorable safety-utility trade-off. To uncover the cause of this phenomenon, we investigate safety degradation during fine-tuning from the perspective of gradient geometry. Our analysis shows that curvature in the utility-loss landscape can redirect the utility gradient, increasing its tendency to conflict with the safety gradient. Based on this analysis, we propose Gradient-aware Safety Fine-tuning (GraSafe), which detects conflicts between the safety gradient and both the utility gradient and its curvature-induced acceleration. Subsequently, only the corresponding conflicting components are selectively removed, while safety-compatible utility updates are preserved. Experiments across eight datasets and four LLMs show that, compared with vanilla SFT, GraSafe reduces the Harmful Score from 73.65% to 7.53% while maintaining comparable downstream task performance, and consistently outperforms existing methods across four tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.