SafetyLens: Multi-Resolution Causal Analysis of Over-Refusal in Large Language Models
Abstract
Safety-aligned language models should carefully mitigate harmful requests. However, these safety mechanisms can also lead to over-refusal (False Positives), where models reject requests that are benign and should have been permitted. There are many reasons why such over-refusals manifest, including disagreements in model usage, usage policies, and, most importantly, the over-sensitivity of existing evaluations. While these evaluations can measure over-refusal behavior, they provide limited insight into the internal computations that give rise to it. To address this issue, we introduce SafetyLens, a mechanistic analysis pipeline that connects behavioral evaluation to internal model calculation. We evaluate SafetyLens on Phi-4, Gemma-4, and Qwen3.6. Using linear classifiers, we analyze the models’ internal representations and identify nodes, including tokens, layers, attention heads, and singular vectors, that are strongly associated with over-refusal. We then perform inference-time ablations on these identified nodes, observing reductions in over-refusal. We further apply this analysis to the input tokens, identifying those whose representations are strongly associated with over-refusal. By replacing these tokens with semantically equivalent alternatives that preserve the original sentence meaning, we reduce over-refusal by 69.03% on Phi-4, 68.72% on Gemma-4, and 59.11% on Qwen3.6. Based on these findings, we reconstruct the over-refusal circuits within each model and identify the components that contribute to the behavior. Intervening on or ablating these circuits at inference time, without modifying the input sentence, reduces over-refusal while preserving the models’ ability to reject harmful queries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.