Safety Lost in Translation: Where Harmful Intent Disappears in Vision-Language Models
Abstract
Vision-language models (VLMs) can refuse a harmful instruction in text yet comply when the same intent is delivered through pixels. Existing work establishes this safety gap across modalities, but does not isolate the computation that separates successful visual understanding from failed refusal. We introduce SIFT (Safety Information Flow Tracing), a paired analysis of the attention message sent from each modality into the shared assistant query state. To separate harmful semantics from refusal itself, refusal control is calibrated on pairs of harmful text prompts with the same underlying goal, one refused and one complied with. We localize the Safety Translation Bottleneck (STB) using a continuous refusal logit gap, with CKA, CCA, and scalar magnitude checks as robustness controls. At the STB, visual messages retain decodable harmful semantics while losing refusal control strength. Grouped edge knockout from source tokens to the query independently peaks in the same depth interval, and bidirectional message swaps on held out examples change the final response rather than only an internal score. The same mechanism motivates SIFT-BRIDGE, a rank 8 translator that repairs only the localized message channel. On a paired bank of 1,680 MM-SafetyBench intents, we obtain 0.93 AUROC for translation failure detection, compared with 0.83 for a paired multilayer probe with matched capacity. Independent replications cover LLaVA-1.5, LLaVA-1.6, Qwen2.5-VL, and InternVL3, while FigStep, VLSBench, JailBreakV-28K, and UnsafeConcepts test transfer beyond rendered text.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.