The Mechanics of Asking: Studying Prompt Requests and Safety Associations
Abstract
Safety evaluations of language models focus heavily on adversarial prompts – jailbreaks, persona assignments, and role-play framings engineered to elicit harmful outputs. However, as language models take on increasingly consequential roles – advising, informing, and acting on behalf of users, safety must hold not only against adversarial attacks but also across the natural ways people phrase a request. If refusal varies across such everyday phrasings, safety becomes vulnerable to its formulation, a particularly consequential weakness in high-stakes domains such as healthcare. In this work, we propose a four-level prompt-based experiment setting grounded in naturalistic patterns, spanning Validation (L0), Delegation (L1), Instructional (L2), and Exploratory (L3) framings of the same underlying action. We characterise refusal across this framework through behavioural evaluations and mechanistic experiments on general-purpose and medically fine-tuned models across different safety benchmarks. Our findings reveal a systematic decline in refusal across levels, with exploratory framing falling by 10–20% on average. Medical fine-tuning further amplifies this vulnerability, even on general-purpose datasets. Mechanistically, we find that models rely on distinct mechanisms to resist framing changes, with some preserving safety through architectural or representational buffering while others lose the internal refusal signal almost entirely. Domain-specific fine-tuning further does not erase the underlying safety mechanism but this vulnerability by altering its sensitivity to how requests are formulated, making exploratory framings harder to refuse. These findings establish consistency across natural variations in request formulation as a critical property of reliable language model safety.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.