acceptodds
Under review as a conference paper at ICLR 2027

When Heads Reveal Safety: Causal Layer Freezing for LLM Fine-Tuning

Abstract

Fine-tuning aligned large language models (LLMs) can weaken their safety behavior, even when the adaptation data are benign. Freezing safety-related components during fine-tuning offers a direct way to preserve alignment, but the component that reveals where safety behavior is localized may not be sufficient as a unit of protection. In particular, freezing individual attention heads can leave other computation pathways in the same layer trainable, whereas freezing a broad contiguous block may unnecessarily restrict task adaptation. We propose Head-Causal Layer Freezing (HeaCaLF), which separates safety localization from protection. HeaCaLF uses activation interventions to estimate the safety contribution of individual attention heads, aggregates these signals by layer, and freezes the selected layers during fine-tuning. Across four aligned LLMs and multiple downstream tasks, HeaCaLF reduces harmful responses relative to standard fine-tuning while retaining competitive task performance. Our granularity analysis further shows that freezing identified heads alone provides limited protection, whereas layer-level freezing offers a more effective safety–utility trade-off.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.