LoRA is All You Need for Refusal-Based Safety Alignment of Reasoning LLMs
Abstract
Reasoning-capable LLMs have achieved major breakthroughs in solving complex problems. However, %recent work shows that acquiring and deploying strong reasoning can introduce significant safety risks. A common mitigation is to apply a secondary safety-alignment phase after reasoning is learned; however, safety alignment often significantly degrades reasoning performance—a phenomenon known as the “Safety Tax”. In this work, we theoretically study the tension in safety fine-tuning: adapting to a new task (safety) while preserving performance on a base task. Our analysis reveals that using a large rank during fine-tuning induces base-task degradation at a rate inversely proportional to the intrinsic dimensionality of the base task. Unlike instruction tuning, reasoning fine-tuning has a large intrinsic dimension. Thus our results suggest applying LoRA for safety fine-tuning on refusal datasets as an effective approach for safety-aligning reasoning models. Despite its simplicity, this recipe achieves safety comparable to full-model alignment while preserving reasoning performance close to the original reasoning-tuned model, across multiple model sizes and architectures, two safety benchmarks, and four reasoning benchmarks spanning mathematics, science, and code generation. Finally, we show that very small-rank LoRA, with as a strong default, achieves the best safety–reasoning trade-off, and that applying LoRA only to MLP modules in middle layers yields superior performance.\looseness=-1
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.