Improving Safety Alignment and Health Safe Completion via Query-Specific Safety Rubrics in OPSD
Abstract
Safety alignment methods have improved resistance to harmful requests, but can still exhibit harmful compliance or unnecessary refusal of benign requests; their ability to provide useful assistance on borderline health queries also remains underexplored (health safe completion). One recent method, OPSA, applies on-policy self-distillation (OPSD) with hard-refusal-oriented teacher privileged information (PI) for harmful queries; in our evaluation, it improves safety but substantially increases over-refusal and reduces health safe completion. Relaxing this PI improves helpfulness but raises harmful compliance, showing that label-level instructions cannot reliably distinguish allowed assistance from prohibited content for each query. We therefore propose Query-Specific Safety Rubrics in OPSD (QSSR-OPSD), which selects and merges rules spanning truthfulness, harmlessness, and helpfulness into general principles, then operationalizes them as query-specific rubrics specifying content to include and avoid. From a probabilistic view of the response distribution, we show that operationalizing general principles into rubrics shifts probability mass toward a set of safe and useful responses when the rubrics reduce penalties on this set more than on other responses, which helps explain our observations both before and after QSSR-OPSD training. Across four backbones, QSSR-OPSD achieves the highest Composite Safety Score (CSS) and health Safe Completion Rate (SCR) in our main results. On Qwen3-8B, it improves CSS from 85.34% to 92.07% and SCR from 36.06% to 74.79% over OPSA, while reducing over-refusal and attack success.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.