Beyond Refusal Rates: How SFT and RL Shape LLM Safety Discernment
Abstract
Safety alignment aims to make large language models helpful without facilitating harm. We investigate how supervised fine-tuning (SFT) and reinforcement learning (RL) shape safety alignment, using group relative policy optimization (GRPO) as a representative RL method. Starting from the same pretrained model and safety prompts, the resulting models achieve similar refusal rates on harmful requests and similar accuracy on math and reasoning benchmarks. Yet GRPO rejects fewer benign requests and more often returns to refusal after an attack forces its response to begin by agreeing to help. These differences concern safety discernment: the ability to answer benign requests and refuse harmful ones despite misleading context. To understand this contrast, we combine training ablations with internal analyses. Rewarding appropriate answers and refusals, rather than harmless content alone, helps GRPO avoid unnecessary refusals. We also find that reinforcing better responses while suppressing worse ones is sufficient to learn recovery after the forced opening. Internal interventions link this recovery to the continued influence of refusal-related information after the forced opening. Together, these findings connect safety discernment to specific training signals and reveal behavioral differences that refusal rates and task accuracy alone fail to capture.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.