Safety Is Discernment, Not Refusal: Teaching Large Reasoning Models to Judge Before They Act
Abstract
Aligned large reasoning models achieve high safety rates on harmful-request benchmarks, yet they substantially over-refuse benign requests that resemble harmful ones. Among the released checkpoints we evaluate on DeepSeek-R1-Distill backbones, every one with a safety rate above 80% answers fewer than half as many of these requests as its base model. This *refusal-shaped safety*, earned by indiscriminate refusal, is invisible to harmful-request-only benchmarks. We therefore argue for evaluating safety as *discernment*: refusing harmful requests while answering benign ones. To teach discernment, we propose **Reason-Judge-Act (RJA)**, a multi-turn reinforcement learning framework. It trains on paired harmful and benign requests that share surface form but differ in intent, so surface cues alone cannot predict the correct behaviour. In each turn, the model reasons about the request, commits to a risk decision in a dedicated judgement slot, and acts on it. The reward penalises over-refusal as heavily as unsafe compliance, and corrective feedback lets the model revise its behaviour in later turns. Across six backbones from two model families, RJA achieves the highest discernment among the evaluated methods while preserving reasoning capability. Notably, with benign non-refusal largely preserved, it not only substantially raises safety on low-safety backbones but also consistently improves it on safer ones. Ablations show that both the corrective turns and the judgement slot contribute to discernment. These results suggest that models should learn to *judge before they act*, not merely to refuse. Data and checkpoints will be released upon acceptance, and code is available at https://anonymous.4open.science/r/RJA.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.