acceptodds
Under review as a conference paper at ICLR 2027

Understanding and Recovering Safety Neurons in VLMs Against Multimodal Jailbreak

Abstract

Vision-language models (VLMs) inherit safety alignment from their language backbones, yet visual inputs can undermine this alignment and enable harmful requests to bypass refusal behavior. We uncover a previously uncharacterized neuron-level phenomenon, termed vision-induced safety-neuron suppression, in which visual inputs may suppress neurons associated with harmful-request refusal and shift their activation patterns toward those induced by benign inputs. Motivated by this finding, we introduce VISAR (Visual Intervention for Safety-neuron Activation Restoration), a lightweight visual intervention framework that uses a differentiable safety probe to supervise a plug-and-play SuffixViT for generating corrective visual suffixes, restoring the suppressed refusal-aligned safety state without modifying the pretrained VLM. Experiments across multiple VLMs and three multimodal jailbreak attacks, compared with four representative defenses, show that VISAR substantially reduces harmful responses while preserving benign utility. Mechanistic analyses further confirm that improved safety is accompanied by restoration of the refusal-aligned safety-neuron activation pattern, demonstrating that visual-side restoration of suppressed safety states provides an effective approach to mitigating multimodal safety degradation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.