acceptodds
Under review as a conference paper at ICLR 2027

AlignLock: Align and Freeze Safety-Critical Layers for Robust LLM Fine-Tuning

Abstract

Safety alignment in large language models (LLMs) is fragile under downstream adaptation: even benign fine-tuning can weaken refusal behaviors and increase susceptibility to jailbreaks. Such degradation arises from the interference between safety alignment and downstream task specialization when they modify shared model parameters. Existing defenses mainly focus on mitigating safety degradation during or after adaptation, which cannot prevent direct overwrite of safety-critical parameters. We propose AlignLock, a proactive layer-level framework for robust safety preservation under downstream adaptation. AlignLock decouples safety alignment and task specialization through a localize-align-lock paradigm: it first identifies a safety-critical layer block via layer-wise sensitivity analysis, performs targeted safety alignment by restricting safety updates to this block, and then locks the same block during downstream fine-tuning to prevent direct overwrite from task gradients. We provide an interference-based analysis showing that locking the safety-critical block eliminates direct overwrite from downstream gradients, while remaining safety drift is limited to indirect interactions through updated task parameters. Experiments across multiple open-weight LLM families and benign downstream domains show consistent safety gains over vanilla fine-tuning with competitive task utility. In the four main model comparisons, AlignLock attains the lowest malicious rates and StrongREJECT scores of the compared methods. It also mitigates malicious fine-tuning under provider-enforced layer locking.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.