acceptodds
Under review as a conference paper at ICLR 2027

Restoring LLM Safety via Sparse Weight Interpolation

Abstract

Supervised fine-tuning of aligned large language models often degrades safety guardrails, even on benign downstream tasks. We identify two distinct failure modes of existing safety-preservation strategies: isolating or restoring safety-critical parameters loses effectiveness under shifted activation contexts, while gradient-based constraints can be circumvented by approximately orthogonal task updates. To avoid such rigid constraints, we formulate safety restoration as a bounded interpolation problem and propose Sparse Weight Interpolation (SWI) to search for a parameter state that jointly recovers safety and preserves task utility. Guided by the aligned base model as a structural prior, SWI optimizes a row-level continuous mask to reactivate refusal behaviors within the fine-tuned parameter space. Empirical evaluations show that SWI restores defense generalization across diverse safety benchmarks with as few as 10 safety demonstrations, while maintaining benign capabilities across multiple model families. We further provide theoretical guarantees for recovery-state compatibility and finite-sample uniform generalization under explicit regularity assumptions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.