Aligning Refusal Behavior with Safety Awareness in Vision-Language Models
Abstract
Vision Language Models (VLMs) have shown remarkable cross-modal reasoning capabilities, yet their safety alignment remains a critical challenge. In this work, we address a fundamental gap in existing VLMs: while these models internally represent harmful content as distinguishable from benign inputs, this safety awareness does not reliably translate into output behavior. Unlike prior approaches that detect and block harmful inputs externally via auxiliary models or additional forward passes, we propose REBASE, a training-free inference-time method that, within a single forward pass, directly bridges the underlying misalignment between the model's internal safety representations and its refusal or compliance behavior through activation steering. Experiments across diverse model architectures and benchmarks demonstrate that our method significantly reduces attack success rates across all settings while maintaining or even improving model utility on benign samples and general vision-language tasks, with less than 2% additional latency. An anonymized version of the code is provided in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.