Safeguards Too Early, Risks Too Late: Toward Safety Alignment Grounded in Multimodal Understanding for LVLMs
Abstract
Large Vision-Language Models (LVLMs) extend LLMs to multimodal queries, yet remain vulnerable to cross-modal compositional jailbreaks (e.g., SIUO, MSSBench), where harmful semantics emerges from image-text interactions rather than either modality alone. This vulnerability reflects a fundamental bottleneck: LVLM safety remains anchored in text-aligned safeguards, leaving safety decisions misaligned with multimodal semantic understanding. To understand how this mismatch arises, we conduct a mechanistic analysis of how LVLMs process jailbreak queries and identify ***Cross-modal Safety Desynchronization (CSD)***: cross-modal harmful semantics forms only after safety-critical window, creating a layer-wise misalignment that lets cross-modal jailbreaks bypass refusal. This raises a key question: *can CSD be bridged by advancing cross-modal fusion into safety-critical window, thereby achieving safety alignment tailored to LVLMs?* Crucially, we find that earlier cross-modal fusion shifts harmful-semantic formation into the safety-critical window, reactivating safety refusal. Building on this, we propose **C**ross-modal **U**nderstanding-guided **R**efusal **E**nhancement (**CURE**), a training-free inference-time defense method that bridges CSD by selectively enhancing cross-modal attention to promote earlier harmful semantic formation, and amplifying safety-critical neuron activations to enable refusal grounded in enhanced multimodal understanding. Experiments on 3 LVLMs across 10 benchmarks show that *CURE* reduces ASR below 3% while preserving general utility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.