When Heads Reveal Safety: Causal Layer Freezing for LLM Fine-Tuning
Abstract
Fine-tuning aligned large language models (LLMs) can weaken their safety behavior, even when the adaptation data are benign. Freezing safety-related components during fine-tuning offers a direct way to preserve alignment, but the component that reveals where safety behavior is localized may not be sufficient as a unit of protection. In particular, freezing individual attention heads can leave other computation pathways in the same layer trainable, whereas freezing a broad contiguous block may unnecessarily restrict task adaptation. We propose Head-Causal Layer Freezing (HeaCaLF), which separates safety localization from protection. HeaCaLF uses activation interventions to estimate the safety contribution of individual attention heads, aggregates these signals by layer, and freezes the selected layers during fine-tuning. Across four aligned LLMs and multiple downstream tasks, HeaCaLF reduces harmful responses relative to standard fine-tuning while retaining competitive task performance. Our granularity analysis further shows that freezing identified heads alone provides limited protection, whereas layer-level freezing offers a more effective safety–utility trade-off.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.