AlignLock: Align and Freeze Safety-Critical Layers for Robust LLM Fine-Tuning
Abstract
Safety alignment in large language models (LLMs) is fragile under downstream adaptation: even benign fine-tuning can weaken refusal behaviors and increase susceptibility to jailbreaks. Such degradation arises from the interference between safety alignment and downstream task specialization when they modify shared model parameters. Existing defenses mainly focus on mitigating safety degradation during or after adaptation, which cannot prevent direct overwrite of safety-critical parameters. We propose AlignLock, a proactive layer-level framework for robust safety preservation under downstream adaptation. AlignLock decouples safety alignment and task specialization through a localize-align-lock paradigm: it first identifies a safety-critical layer block via layer-wise sensitivity analysis, performs targeted safety alignment by restricting safety updates to this block, and then locks the same block during downstream fine-tuning to prevent direct overwrite from task gradients. We provide an interference-based analysis showing that locking the safety-critical block eliminates direct overwrite from downstream gradients, while remaining safety drift is limited to indirect interactions through updated task parameters. Experiments across multiple open-weight LLM families and benign downstream domains show consistent safety gains over vanilla fine-tuning with competitive task utility. In the four main model comparisons, AlignLock attains the lowest malicious rates and StrongREJECT scores of the compared methods. It also mitigates malicious fine-tuning under provider-enforced layer locking.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.