acceptodds
Under review as a conference paper at ICLR 2027

Harm Recognition to Refusal Gap: Why Small LLMs are unsafe?

Abstract

The growing deployment of small language models makes their vulnerability to jailbreaks an important safety concern. A common belief is that small models are less robust to jailbreaks because they have weaker representation, reasoning, and instruction-following abilities, and thus fail to recognize disguised harmful intent. However, we find that small models can still clearly separate harmful from benign inputs in their internal representations. Their failures arise because this harmfulness-related signal does not reliably translate into refusal behavior. Using existing harmfulness and refusal directions, we show that prompts with strong harmfulness-related activations can still elicit compliant or harmful responses. This suggests a bottleneck in the internal conversion from harmfulness recognition to refusal execution, which we call the recognition-to-refusal gap. Across model scales, safety differences are more closely associated with the coupling between these two signals than with refusal strength alone. Motivated by this finding, we propose Coupling-Aware Direct Preference Optimization (CA-DPO). CA-DPO focuses training on jailbreak cases where harmfulness recognition does not lead to refusal, while retaining benign examples to preserve helpfulness and limit over-refusal. Experimental results show that CA-DPO improves model safety while mitigating over-refusal, suggesting that optimizing the recognition-to-refusal pathway can improve safety while preserving useful behavior in safety-critical deployment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.