Language Model Refusal Changes with Users' Stated Chance of Being in a Position to Act
Abstract
Safety-aligned language models are trained to refuse harmful requests, but it is unclear whether refusal changes with the user's stated chance of being in a position to act on the assistance. We test this by appending one sentence to each request stating this chance, varying it from to while holding the request fixed. We evaluate 16 models on 84 harmful and hard-benign requests from OR-Bench and measure refusal using two independent classifiers. We find that refusal increases with the stated chance under both measures, by percentage points under the judge classifier and points under the rule classifier. The two classifiers give different point estimates, and blinded human adjudication does not favor either measure. We also find that refusal increases with the stated chance on the benign requests, with no evidence that this increase is larger on harmful than benign requests. Overall, our findings show that refusal is sensitive to the stated chance that a user can act, but this sensitivity alone does not show that refusal is better targeted to harmful requests.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.