LoRA Safety Patches: Capability-Preserving Refusal Tuning with Safety Data Alone
Abstract
Safety alignment of large language models (LLMs) often improves resistance to harmful instructions and jailbreak attacks but degrades general capabilities. Existing approaches mitigate this trade-off by mixing safety and general-purpose data, making alignment sensitive to data composition. In this work, we empirically and mechanistically investigate whether standard LoRA can improve safety during safety alignment without sacrificing pretrained capabilities. Across multiple LLM families, LoRA-based refusal tuning using only safety data substantially improves safety while largely preserving general capabilities. Compared with full-parameter tuning, it also supports continual safety alignment with negligible utility loss and exhibits less safety regression under continuous benign fine-tuning. To better understand this behavior, we conduct empirical analyses of parameter updates, hidden-state shifts, and subspace similarity. Our results show that LoRA-based safety updates introduce substantially lower interference with the model’s intrinsic transformations than full-parameter tuning, suggesting a form of subspace decoupling that enables lightweight and modular safety patching for LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.