acceptodds
Under review as a conference paper at ICLR 2027

SAFETY ALIGNMENT IS CONTEXT-BLIND: MODELS DETECT DISTRESS BUT DO NOT ACT ON IT

Abstract

Large language models can recognise emotional distress in user context, yet this information may fail to influence their safety decisions. We identify and characterise this failure mode as context-blind responding, in which models detect distress but fail to integrate it with the risk of a subsequent request. Using controlled counterfactual scenarios, we evaluate two open-weight models (Llama 3.1 8B and Qwen3 4B) and two proprietary models (GPT-5.4 and Gemini 3.6) and find that the two open-weight models answer roughly three-quarters of distress-conditioned risky requests that should be declined, while proprietary models continue to answer a substantial fraction. Mechanistic analysis reveals that distress is linearly decodable from the residual stream as early as the first transformer layer, yet remains weakly aligned with the refusal direction. Causal tracing further shows that the effect of distress on the model's internal refusal signal emerges only in a late network region at approximately 72% of network depth, indicating a gap between contextual representation and safety computation. Motivated by this finding, we introduce a conjunctive gate that combines risk detection with the refusal direction at this integration region, substantially improving selective refusal while preserving helpfulness. These results suggest that an important class of safety failures arises not from missing information, but from failures to integrate information already represented within the model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.