Safety Guards Judge Harmfulness, Not Policy Violation
Abstract
Policy-conditioned safety guards are designed to judge a prompt against the policies enabled for them, not merely by whether it is harmful. Across the evaluated open guards, verdicts on harmful prompts rarely flip when the enabled policy stops covering them, even when the guard correctly identifies the policy they violate. Linear probes show that hidden states encode this input–policy relation, yet verdicts mainly follow an estimated harmfulness direction, and in several guards intervening on the relation leaves them largely unchanged. To reconnect this relation to the verdict, we introduce *policy-gated harmfulness routing*, which uses it to gate an adjustment along the harmfulness direction without updating the guard's weights. Routing improves *paired accuracy* on the evaluated guards, with some loss in target-policy detection and no increase in benign on-topic false-positive rates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.