ResInject: State-Dependent Residual Adaptation for Joint Safety and Over-Refusal Control
Abstract
Safety alignment requires Large Language Models (LLMs) to accurately separate harmful requests that must be declined from benign queries that should be served. Because this safety boundary is learned implicitly rather than specified explicitly, benign prompts sharing lexical similarities with harmful ones often fall on the wrong side, leading to over-refusal. Reducing over-refusal without increasing unsafe generations is notoriously difficult, as both behaviors are governed by the same decision boundary. Prior methods typically rely on external preference datasets or decoding-time representation steering to shift this boundary. In this work, we leverage the inherent stochasticity of the model at the safety boundary. We observe that for borderline prompts, a frozen aligned model naturally generates a mixture of refusals, unsafe responses, and safe compliances. We formalize these self-generated outputs as on-policy preference pairs and propose RESINJECT, a hidden-state intervention technique that trains a low-rank residual stream adapter on these contrastive pairs. By excluding explicit refusal targets, our reference-adjusted preference objective indirectly suppresses over-refusal while preserving the rejection of genuinely harmful requests. Extensive experiments on Llama-3, Qwen-2.5, and OLMoE demonstrate that our method significantly reduces over-refusal on safety boundary prompts and external benchmarks without compromising safety against harmful queries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.