SAFETY ALIGNMENT IS CONTEXT-BLIND: MODELS DETECT DISTRESS BUT DO NOT ACT ON IT
Abstract
Large language models can recognise emotional distress in user context, yet this information may fail to influence their safety decisions. We identify and characterise this failure mode as context-blind responding, in which models detect distress but fail to integrate it with the risk of a subsequent request. Using controlled counterfactual scenarios, we evaluate two open-weight models (Llama 3.1 8B and Qwen3 4B) and two proprietary models (GPT-5.4 and Gemini 3.6) and find that the two open-weight models answer roughly three-quarters of distress-conditioned risky requests that should be declined, while proprietary models continue to answer a substantial fraction. Mechanistic analysis reveals that distress is linearly decodable from the residual stream as early as the first transformer layer, yet remains weakly aligned with the refusal direction. Causal tracing further shows that the effect of distress on the model's internal refusal signal emerges only in a late network region at approximately 72% of network depth, indicating a gap between contextual representation and safety computation. Motivated by this finding, we introduce a conjunctive gate that combines risk detection with the refusal direction at this integration region, substantially improving selective refusal while preserving helpfulness. These results suggest that an important class of safety failures arises not from missing information, but from failures to integrate information already represented within the model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.