acceptodds
Under review as a conference paper at ICLR 2027

Same Harm, Different Responses: Reliance on Surface-Level Cues in Safety-Aligned LLMs

Abstract

Safety alignment enables large language models (LLMs) to identify and refuse policy-violating requests, serving as a key approach to reducing the risk of harmful content generation. However, the continued emergence of jailbreak attacks demonstrates that these safeguards can be bypassed through low-cost manipulations, such as weakening or restructuring risk cues in harmful requests. To better understand and mitigate this vulnerability, we investigate how the visibility of surface-level risk cues affects safety behavior and reveal a behavioral pattern that has not yet been systematically characterized. Specifically, simply adding explicit risk cues, such as pressure or temptation, before an unchanged harmful request substantially reduces attack success and response harmfulness. To validate this pattern, we construct paired prompts that preserve the underlying harmful intent while varying risk-cue visibility and conduct controlled experiments across multiple open-source LLMs. The results consistently demonstrate this effect across different risk cues and model settings, revealing a systematic pattern consistent with model reliance on the visibility of surface-level risk cues and suggesting that safety-aligned models may not yet have established robust intent-level safety boundaries. We further analyze how such reliance may emerge during safety alignment, providing a theoretical account of one potential mechanism behind the persistent effectiveness of low-cost jailbreak attacks. Building on this analysis, we improve safety alignment methods by encouraging consistent safety behavior across requests that share the same harmful intent but differ in risk-cue visibility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.