acceptodds
Under review as a conference paper at ICLR 2027

LoRA Safety Patches: Capability-Preserving Refusal Tuning with Safety Data Alone

Abstract

Safety alignment of large language models (LLMs) often improves resistance to harmful instructions and jailbreak attacks but degrades general capabilities. Existing approaches mitigate this trade-off by mixing safety and general-purpose data, making alignment sensitive to data composition. In this work, we empirically and mechanistically investigate whether standard LoRA can improve safety during safety alignment without sacrificing pretrained capabilities. Across multiple LLM families, LoRA-based refusal tuning using only safety data substantially improves safety while largely preserving general capabilities. Compared with full-parameter tuning, it also supports continual safety alignment with negligible utility loss and exhibits less safety regression under continuous benign fine-tuning. To better understand this behavior, we conduct empirical analyses of parameter updates, hidden-state shifts, and subspace similarity. Our results show that LoRA-based safety updates introduce substantially lower interference with the model’s intrinsic transformations than full-parameter tuning, suggesting a form of subspace decoupling that enables lightweight and modular safety patching for LLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.