acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Harmful Fine-Tuning via Task-Aware Directional Safety Repair

Abstract

Fine-tuning large language models on unvetted data can compromise their safety alignment, particularly when the training data contain harmful examples. However, preserving safety throughout harmful fine-tuning without hindering downstream task adaptation remains challenging. To address this challenge, we propose Task-Aware Directional Safety Repair (TDSR), a defense that integrates periodic safety repair into downstream fine-tuning. During each safety repair, TDSR uses gradients of the safety and downstream task losses to identify candidate directions along which parameter updates strongly affect safety losses while causing relatively small changes in task losses. It then selects the candidate most closely aligned with the direction of steepest decrease in the current safety loss and projects the safety descent gradient onto the selected candidate direction to update the model. Extensive experiments across various models and downstream tasks show that TDSR preserves safety alignment while maintaining task performance, outperforming existing baselines. For example, on Mistral-7B-v0.3, TDSR reduces the mean harmful-response rate by 15.62 percentage points and improves mean task performance by 11.30 percentage points across three tasks, relative to the strongest baselines for safety and downstream task performance, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.