Beyond Parameter Space: Closed-Form Cross-Trajectory Repair After Harmful Fine-Tuning
Abstract
Safety-aligned language models are routinely adapted after deployment to improve downstream task performance, yet the fine-tuning process may be vulnerable to a small fraction of harmful examples that compromise their safety alignment. Existing safety repair methods intervene through diverse mechanisms, including parameter or update modification, model merging, and additional optimization. However, harmful fine-tuning can also change the internal representations propagated through the model, so the aligned reference and attacked models may reach the same module with different representations even for the same input. Motivated by this observation, we propose CCTR (Closed-Form Cross-Trajectory Repair), a post-hoc repair method that fits the full cross-trajectory output residual on representations produced by the attacked model while constraining output drift on task-agnostic utility activations. For each target linear module, these two objectives form a regularized least-squares problem with a closed-form solution, enabling repair without gradient-based optimization or downstream-task supervision. Compared with the strongest state-of-the-art safety repair method, CCTR lowers the average Harmful Score from 2.30% to 0.38% across varying harmful-data ratios. In experiments on different downstream tasks, it lowers the score from 2.85% to 0.77% and improves average task accuracy by 3.53%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.